TapVid
    API & MCPPricingBlogAbout
    Blog›AI Video Benchmark: Accuracy vs Aesthetics
    Back to Blog

    AI Video Benchmark: Accuracy vs Aesthetics

    See what 6,760 blind ratings reveal about factual coverage, viewer comprehension, visual quality, and aesthetics in AI-generated video.

    Research
    Derek LawDerek LawSeptember 7, 2026 · 29 min readSep 7, 2026 · 29 min readDiscord
    Derek LawDerek LawCTO & Co-founder, TapVid

    Connect with the author, meet other video creators, and watch hands-on tutorials.

    Join our Discord
    September 7, 202629 min read
    AI video benchmark comparing accuracy, comprehension, visual quality, and aesthetics across seven systems
    Summarize with6 assistants
    ChatGPTPerplexityTapVidvideoClaudeGeminiGrok
    Create videos from your AI agentConnect TapVid API & MCP→

    In this article

    1. 01AI video benchmark verdict: first on the content that matters — balanced everywhere else
    2. 02Built for videos that must inform, not entertain
    3. 03Informative video is a trade-off frontier — and TapVid holds the balance point
    4. 04Getting facts on screen is the binding constraint — and TapVid does it best
    5. 05Viewers learn as much from TapVid as from any system in the field
    6. 06Statistically tied with the leader on three of four visual dimensions
    7. 07Human annotations validated by VBench industry-standard automated metric
    8. 08Two targets: Marketing content, and aesthetic appeal without an informative trade
    9. 09Frozen protocol, seeded statistics, every result in full
    10. 10Try TapVid
    11. 11References and citation
    12. 12About the author
    Summarize withAPI & MCP →
    ChatGPTPerplexityTapVidClaudeGeminiGrok

    TL;DR

    On information-dense video, TapVid leads the field.

    This AI video benchmark is built for video that must inform — instructional, editorial, explanatory — not entertain. On the information-dense scenes it targets, TapVid takes the outright Factual Coverage lead, ties for first on TeachQuiz Information Delivery, and tops Overall Quality across seven systems. The independent, industry-standard automated metric, VBench, reproduces the human panel’s preference, confirming the eval’s rigor.

    01

    AI video benchmark verdict: first on the content that matters — balanced everywhere else

    Information-dense scenes — professional, news, how-to, education — are the benchmark’s target domain, and they are where TapVid wins. Across all 25 scenes, TapVid is the clear 2nd of 7 on Informative Quality and the only code renderer strong on both informative axes .

    Overall Quality — informative + aesthetic 75.2 T #1 of 7 on information-dense work — the only system high on every axis at once.

    Aesthetic Preference vs Fable 5, the informative leader 58.7% Preferred over the field’s top informative system in side-by-side viewing — and significantly does-not-lose.

    Factual Coverage — information-dense scenes 83.6% The outright field lead. Significantly ahead of 5 of 6 competitors at getting required facts on screen.

    TeachQuiz Information Delivery — information-dense scenes 84.5% Tied for first. Viewers answer comprehension questions as well after TapVid as after any system in the field.

    OVERALL QUALITY — INFORMATIVE + AESTHETIC · INFORMATION-DENSE SCENES

    *(Content Fidelity + Visual Quality + Aesthetic Preference) / 3 · field-relative T-score (70 = field mean, 10 = 1 SD)*

    OVERALL QUALITY — INFORMATIVE + AESTHETIC · INFORMATION-DENSE SCENES
    OVERALL QUALITY — INFORMATIVE + AESTHETIC · INFORMATION-DENSE SCENES

    Both frames are reported everywhere. Headline claims use the information-dense frame (UC1–UC4) the benchmark targets — shown here, where TapVid tops Overall Quality; on all 25 scenes TapVid is 2nd of 7 on Informative Quality (T 78 vs Fable 5’s 85). Full scorecards with confidence intervals are in the appendix .

    THE BENCHMARK

    02

    Built for videos that must inform, not entertain

    Each of five use cases draws on real-world source material and is evaluated on five intent-driven scenes: the brief specifies the facts a viewer should walk away with, and human annotators measure whether they do. The target is high-information-density, clear-intent video — instructional, editorial, explanatory — not aesthetic or entertainment content.

    THE FIELD — SEVEN SYSTEMS, TWO PARADIGMS

    SystemHarness / modelFamily
    TapVidcode-rendered videocode-render
    Fable 5Claude Code · Fable 5code-render
    Code-render HClaude Code · Sonnet 5code-render
    Code-render RCodex · GPT-5.5code-render
    Pixel-diffusion Lofficial studio · Pro mode, 1080ppixel-diffusion
    SeedanceSeedance 2.0 · Jimeng · VIP mode, 720ppixel-diffusion
    VeoVeo 3.1 · Google Flow · Quality mode, 720ppixel-diffusion

    All systems accessed 2026 Jul 6–12 : locally-run tools frozen at their latest release as of Jul 6, hosted systems at the latest flagship model on their official platform — every system at its maximum quality setting, 16:9. Exact version pins with sources are in appendix A .

    FIVE USE CASES, REAL-WORLD SOURCE MATERIAL

    Use caseSource material
    Professional Enterprise / self-learningGDPval tasks — Professional & Technical Services, Government, Manufacturing, Finance
    News Editorial2025–26 covers: Economist, FT, Bloomberg, Nature, Science
    How-To LifestyleTop #HowTo YouTube — craft/DIY, art, fitness, mindset, puzzles
    Explainer EducationMost-viewed “AP Exam explained” — history, psych, calculus, biology, geography
    Marketing Product / salesNASDAQ most-valuable-company launch videos across sectors

    Sampling rules fixed before generation. Each use case draws its five scenes by a predefined rule — cluster the source pool, randomly sample one per cluster (top industries, publications, content categories, AP subjects, NASDAQ sectors). Every brief is a fully-detailed video-generation prompt carrying the complete information payload, authored and frozen before any system generated anything ; its required-fact checklist (168 facts across the benchmark) and comprehension questions (149) were authored at the same time and human-reviewed. A complete worked example is in appendix B .

    HOW THE FIELD WAS RUN — ONE SHOT, SAME BRIEF, TOP SETTINGS

    • Identical inputs: every system received the same brief verbatim — zero per-system prompt engineering.
    • Best-of-1: one generation per scene, never curated; hard failures retried until success (at most 3 retries observed).
    • Durations as generated: target lengths set by the brief; each system’s actual output used as-is (system averages 1.1–4.2 min).
    • The 173/175 accounting: Fable 5’s safety rules declined 2 of its 25 briefs. Rather than substitute its Opus fallback, we scored Fable 5 only on its own 23 generations — no competitor was quietly weakened, and missing cells are dropped from denominators, never imputed.

    THE ANNOTATOR PANEL — INDEPENDENT, BLIND, FULLY CROSSED

    • 12 independent annotators , outsourced freelancers unaffiliated with any system’s team — each holding a master’s degree or above from a QS World Ranking top-50 university, with certified English competence.
    • Blind throughout: videos presented unlabeled in randomized order; annotators were never told which system produced a video.
    • Fully crossed: all 12 annotators rated all 7 systems, two raters per video — a strict or lenient rater moves every system equally and cannot bias a comparison.
    • Deliberate viewing, not click-work: ~20 minutes of annotation per video of under 5 minutes, paid at 4× local minimum wage. Annotated Jul 13–15, 2026.

    Metrics grounded in published methods, measured by people

    Every video is scored by a fully-crossed annotator panel — every annotator rated every system, so a strict or lenient rater moves all systems equally and cannot bias a comparison. The panel produced 6,760 human annotations across five stations grounded in published methods — three of them peer-reviewed at ICML, WACV, and IEEE TPAMI / CVPR — and every video is *also* scored by the VBench automated model: >3.7 million frame-level model evaluations .

    • TeachQuiz Information Delivery (Code2Video, ICML 2026 · arXiv:2510.01174) — how much of the intended information a viewer actually acquires.
    • Factual Coverage — of the facts the brief requires, how many appear on screen, verbatim and legible.
    • T2VTextBench Text Accuracy (arXiv:2505.04946) — correctness of on-screen titles, numbers and labels.
    • T2VWorldBench ( WACV 2026 · arXiv:2507.18107) — four quality dimensions: Video Quality, Realism, Relevance, Consistency.
    • Aesthetic Preference — arena-style, side-by-side comparisons against TapVid.
    • VBench (VBench++, IEEE TPAMI 2025 · original VBench, CVPR 2024 Highlight · the model behind the public HuggingFace leaderboard) — industry-standard automated video-quality suite, >3.7 million frame-level model evaluations .

    VBench is an independent, machine-based cross-check of our human annotations : its Frame Quality evaluation reproduces our human preference ranking, so the two instruments corroborate each other ( Finding 05 ).

    The rankings are robust to how you compute them: three independent estimators — raw rates, leave-one-out information gain, and mixed-effects logistic models — produce the identical ordering, and reliability-weighted and PCA variants leave it unchanged. Because Marketing behaves as a different task family from the four information-dense use cases, every result is reported in two frames : all 25 scenes, and the information-dense frame (Professional, News, How-To, Explainer). Full tables for both are in the appendix .

    FINDING 01 · THE FULL-FIELD FRAME

    03

    Informative video is a trade-off frontier — and TapVid holds the balance point

    WIN The field is a trade-off frontier: on Informative Quality the order is Fable 5 > TapVid > everyone else; on Aesthetic Preference the order inverts exactly. TapVid is the only code renderer high on both informative axes — Content Fidelity 77, Visual Quality 79, 2nd of 7 on each — the balance point of the frontier.

    Aesthetic Preference is measured arena-style — two videos of the same brief, side by side — and captures something the informative rubrics do not: its correlation with every informative metric is negative-to-zero (−0.22 to −0.01). Code renderers buy informative quality; pixel-diffusion buys aesthetic appeal. The frontier below places all seven systems on that trade-off, in both frames.

    THE INFORMATIVE–AESTHETIC FRONTIER

    *x: Informative Quality · y: Aesthetic Preference · field-relative T-scores (70 = field mean) · information-dense scenes*

    THE INFORMATIVE–AESTHETIC FRONTIER
    THE INFORMATIVE–AESTHETIC FRONTIER

    TapVid is the only system in the green quadrant — highly informative and aesthetically preferred — in both frames. The frontier runs downhill: Informative Quality orders Fable 5 > TapVid > the rest, and Aesthetic Preference inverts it end-for-end. On information-dense scenes TapVid moves to the frontier’s elbow (Informative Quality 82, within 3.5 T of the leader) while the pixel-diffusion aesthetic edge shrinks to parity; on all 25 scenes Seedance and Pixel-diffusion L sit above the preference range shown. Aesthetic Preference is reported beside Informative Quality, never averaged into it.

    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, ALL 25 SCENES

    *share of side-by-side comparisons TapVid wins (ties count half) · ×100*

    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, ALL 25 SCENES
    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, ALL 25 SCENES

    Above 50 = viewers prefer TapVid. Significance from the mixed-logit non-loss / decisive-win tests; full CIs and the non-loss frame are in the appendix .

    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, INFORMATION-DENSE SCENES

    *share of side-by-side comparisons TapVid wins (ties count half) · ×100*

    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, INFORMATION-DENSE SCENES
    AESTHETIC WIN-RATE OF TAPVID — SIDE-BY-SIDE VIEWING, INFORMATION-DENSE SCENES

    On the benchmark’s target domain the aesthetic deficits shrink to parity — only Seedance remains significantly preferred on decisive comparisons.

    WIN Shown side-by-side against Fable 5 — the field’s informative leader — viewers prefer TapVid. Win-rate 58.7%, and TapVid significantly does-not-lose (P = 0.72). The system that beats TapVid on quality composites loses to it in front of viewers.

    WIN On information-dense work, TapVid concedes nothing on aesthetics while leading content. Overall non-loss rises to 53.0% — parity — and no pixel-diffusion system holds a significant non-loss advantage there. Against Google’s Veo: ahead on Informative Quality (78 vs 65), significantly ahead on coverage, and a preference toss-up.

    LIMITATION Across all 25 scenes, aesthetic preference is a real weakness. TapVid’s overall win-rate is 40.7% — significantly below the 50% no-preference line — and it is significantly less preferred than Pixel-diffusion L and Seedance, the bottom of the informative ranking. The inversion is the paradigm trade-off, concentrated in Marketing; the close-out plan addresses it directly.

    FINDING 02 · FACTUAL COVERAGE

    04

    Getting facts on screen is the binding constraint — and TapVid does it best

    WIN On information-dense scenes TapVid puts 83.6% of required facts on screen — the outright field lead, significantly ahead of 5 of 6 competitors. Across all 25 scenes it is statistically tied with the leader Fable 5 and significantly ahead of every pixel-diffusion system.

    Coverage is the benchmark’s most discriminating metric: the field splits into two clean tiers, with the four code renderers (74–79% of facts on screen) sitting well clear of the three pixel-diffusion systems (54–61%).

    FACTUAL COVERAGE — SHARE OF REQUIRED FACTS ON SCREEN

    *mean · 95% bootstrap CI over scenes · information-dense scenes · ×100*

    FACTUAL COVERAGE — SHARE OF REQUIRED FACTS ON SCREEN
    FACTUAL COVERAGE — SHARE OF REQUIRED FACTS ON SCREEN

    Info-dense frame: TapVid 83.6 [77.3–89.0] leads outright ; the gap to every system except Fable 5 is statistically significant (mixed-logit 95% CI excludes 0). All-25 frame: Fable 5 79.0, TapVid 77.3 — a statistical tie — with Seedance, Veo and Pixel-diffusion L significantly behind in both frames.

    COVERAGE DRIVES DELIVERY — ONE POINT PER VIDEO, 173 VIDEOS

    *x: Factual Coverage · y: TeachQuiz Information Delivery (correct-rate) · ×100*

    COVERAGE DRIVES DELIVERY — ONE POINT PER VIDEO, 173 VIDEOS
    COVERAGE DRIVES DELIVERY — ONE POINT PER VIDEO, 173 VIDEOS

    WIN Once a fact is on screen, viewers acquire it ~95% of the time. Coverage and delivery correlate at r = 0.70 across 173 videos, and the fitted line runs from a 36% no-fact floor to a 96% on-screen ceiling. Comprehension is nearly a solved step — the game is coverage, and TapVid plays it best.

    COVERAGE BY USE CASE — MODEL-PREDICTED P(FACT ON SCREEN)

    *mixed-logit predicted probability per use case · ×100*

    COVERAGE BY USE CASE — MODEL-PREDICTED P(FACT ON SCREEN)
    COVERAGE BY USE CASE — MODEL-PREDICTED P(FACT ON SCREEN)

    WIN TapVid leads or ties for the coverage lead in four of five use cases. It leads News (81.9) and How-To (86.7) outright, and on Professional it is significantly ahead of four competitors — including the overall leader Fable 5.

    LIMITATION Marketing is the exception. TapVid’s coverage drops to 53.8% there — bottom of the code renderers, behind Fable 5, Code-render R and Code-render H. The same gap shapes every metric’s Marketing column; the close-out plan treats it as one problem.

    FINDING 03 · TEACHQUIZ INFORMATION DELIVERY

    05

    Viewers learn as much from TapVid as from any system in the field

    WIN On information-dense scenes TapVid ties for first on TeachQuiz Information Delivery — viewers answer 84.5% of comprehension questions correctly — and significantly out-delivers three systems (Code-render H, Seedance, Pixel-diffusion L). Across all 25 scenes the top five are statistically inseparable; TapVid is significantly ahead of Pixel-diffusion L and Seedance.

    Delivery follows the TeachQuiz protocol: after watching, annotators answer multiple-choice questions about the facts the brief required, and each video is scored against a per-question field baseline. Three independent estimators — raw correct-rate, leave-one-out information gain, and a mixed-effects logit — agree on the ordering.

    TEACHQUIZ INFORMATION DELIVERY — COMPREHENSION CORRECT-RATE

    *raw correct-rate · significance vs TapVid from mixed-effects logit · information-dense scenes · ×100*

    TEACHQUIZ INFORMATION DELIVERY — COMPREHENSION CORRECT-RATE
    TEACHQUIZ INFORMATION DELIVERY — COMPREHENSION CORRECT-RATE

    TapVid is top-tier in both frames. Full per-use-case win/tie/loss tables are in the appendix .

    DELIVERY BY USE CASE — COMPREHENSION CORRECT-RATE

    *raw correct-rate per use case · bars with 95% bootstrap CI over the use case’s scenes · ×100*

    DELIVERY BY USE CASE — COMPREHENSION CORRECT-RATE
    DELIVERY BY USE CASE — COMPREHENSION CORRECT-RATE

    TapVid is top-tier on News (88.3) and How-To (89.7) — per the mixed-logit inference it significantly out-delivers Code-render R, Pixel-diffusion L and Seedance on News and Veo on How-To, losing to none in either. Professional and Explainer are mid-field; Marketing is the gap the close-out plan targets. Each use case rests on 5 scenes, so the CIs are wide; significance calls come from the per-use-case mixed logit in the appendix .

    T2VTEXTBENCH TEXT ACCURACY — ON-SCREEN TITLES, NUMBERS AND LABELS

    *four-tier grade mapped to [0,1] · 95% bootstrap CI · information-dense scenes · ×100*

    T2VTEXTBENCH TEXT ACCURACY — ON-SCREEN TITLES, NUMBERS AND LABELS
    T2VTEXTBENCH TEXT ACCURACY — ON-SCREEN TITLES, NUMBERS AND LABELS

    WIN TapVid renders on-screen text best on information-dense scenes. Across all 25 scenes the top of the field is statistically inseparable.

    FINDING 04 · VISUAL QUALITY

    06

    Statistically tied with the leader on three of four visual dimensions

    WIN On the T2VWorldBench human evaluation, TapVid is statistically tied with the leader Fable 5 on Realism, Relevance and Consistency — in both frames — and a clear 2nd of 7 on the four-dimension composite (86.4), significantly ahead of Seedance and Pixel-diffusion L. Whatever TapVid renders, it renders cleanly.

    The four dimensions are rated 1–5 by the crossed annotator panel against per-dimension rubrics. The one significant visual gap to Fable 5 is Video Quality (86.0 vs 90.0); the other three dimensions are statistical ties.

    T2VWORLDBENCH — FOUR DIMENSIONS, ALL 25 SCENES

    *mean grade · 95% bootstrap CI over scenes · ×100 · * = significant vs TapVid*

    T2VWORLDBENCH — FOUR DIMENSIONS, ALL 25 SCENES
    T2VWORLDBENCH — FOUR DIMENSIONS, ALL 25 SCENES

    TapVid is 2nd on Quality, Realism and Consistency, 3rd on Relevance (behind Fable 5 and Code-render R). Relevance — “does it say the right things” — behaves as a content dimension, and is analyzed with the content metrics above.

    WIN Where raters fully agree, TapVid rates highest. On the exact-agreement subset TapVid’s Video Quality rises to 0.97, overtaking Fable 5 — TapVid’s clear wins are unambiguous.

    FINDING 05 · INDEPENDENT CORROBORATION

    07

    Human annotations validated by VBench industry-standard automated metric

    CHECK Run blind on the same 25 scenes, VBench — the automated model behind the public HuggingFace leaderboard — scores every video independently of the human panel. It orders the systems the same way our annotators’ side-by-side preference does. Spearman ρ = +0.89 between VBench Frame Quality and human preference; +0.87 leaving any one system out. The human preference axis is not an artifact of our raters — an independent, paper-grounded machine metric lands in the same order.

    A widely-used automated model, computed with no knowledge of the human labels, agrees with the human panel on how the systems rank — evidence that the eval captures a real, reproducible signal rather than rater idiosyncrasy.

    AUTOMATED AND HUMAN INSTRUMENTS AGREE — EVERY VIDEO SHOWN

    *x: VBench Frame Quality per video (automated) · y: each system’s human win-rate · 141 system×scene videos, coloured by system · ×100*

    AUTOMATED AND HUMAN INSTRUMENTS AGREE — EVERY VIDEO SHOWN
    AUTOMATED AND HUMAN INSTRUMENTS AGREE — EVERY VIDEO SHOWN

    Each small dot is one video’s automated Frame Quality, coloured by system and placed at that system’s human win-rate; the large dots are the system means. The coloured clouds stack in win-rate order and their centres climb left-to-right — automated frame quality and human preference rank the systems together ( Spearman ρ = +0.89 ). Systems are left unnamed: the point is the agreement of the two instruments, not any one system’s position.

    SRC VBench (VBench++, IEEE TPAMI 2025 · original VBench, CVPR 2024 Highlight) is the standard automated video-quality benchmark behind the public HuggingFace leaderboard, scoring six dimensions per clip. It was run with the official VBench-Long code path (commit e946a8d) on dedicated GPUs on Jul 13–14 — completed before any comparison against the human labels existed. The Frame Quality head (Aesthetic + Imaging Quality) is the aesthetic cross-check used here, computed quality-when-it-works so a failed generation does not distort a system’s score. Every composite was recomputed from the raw per-scene scores and matches the source report to the decimal; the full run protocol is in appendix A.

    THE NEXT WINS TO CLOSE

    08

    Two targets: Marketing content, and aesthetic appeal without an informative trade

    Both gaps are measured, localized, and separable from what already works. Neither calls for touching the coverage-and-delivery engine that wins the information-dense domain.

    TAPVID BY USE CASE — THE MARKETING GAP IS CONTENT, NOT RENDERING

    *Content Fidelity vs Visual Quality composite · field-relative T-score · per use case*

    TAPVID BY USE CASE — THE MARKETING GAP IS CONTENT, NOT RENDERING
    TAPVID BY USE CASE — THE MARKETING GAP IS CONTENT, NOT RENDERING

    On Marketing, TapVid’s Visual Quality holds at 75 while Content Fidelity drops to 51 — the videos render cleanly but miss required facts. Everywhere else the two axes move together at 68–93.

    THE AESTHETIC PROGRAM — MOVE UP THE AESTHETIC AXIS, HOLD THE INFORMATIVE LEAD

    *x: Informative Quality · y: Aesthetic Preference · T-scores · all 25 scenes*

    THE AESTHETIC PROGRAM — MOVE UP THE AESTHETIC AXIS, HOLD THE INFORMATIVE LEAD
    THE AESTHETIC PROGRAM — MOVE UP THE AESTHETIC AXIS, HOLD THE INFORMATIVE LEAD

    Because preference is orthogonal to every quality metric, the move is straight up — raising motion, polish and appeal shifts TapVid toward the frontier’s empty top-right corner without spending any of its informative lead.

    NEXT Win to close #1 — Marketing content. The collapse is fact delivery in marketing scenes, not rendering — a targeted extension of the coverage engine that already leads four use cases, not a rebuild. Closing it converts TapVid’s information-dense coverage lead into a full-field one: excluding Marketing, TapVid already tops five of six systems on coverage outright.

    NEXT Win to close #2 — aesthetic preference vs pixel-diffusion, without giving back the informative lead. Preference is empirically orthogonal to every quality metric, so closing it is its own program — motion, polish, appeal — not a trade forced against coverage or delivery. The information-dense frame shows the ceiling: where content is dense, TapVid already reaches preference parity while holding the content lead.

    APPENDIX

    09

    Frozen protocol, seeded statistics, every result in full

    System versions are pinned to the day with sources; sampling rules and fact checklists were fixed before any video existed; the annotator panel was blind and fully crossed (A–B). Three independent estimators agree on every ordering, all randomness is seeded, and an independent re-run reproduced every artifact (C–E). Complete leaderboards with confidence intervals for every metric, in both frames (F–K). Δ columns are mixed-logit log-odds contrasts vs TapVid unless noted; WIN / LOSS = the 95% interval excludes zero, from TapVid’s perspective; TIE = it does not.

    KEY System naming. The body of this report refers to three systems by anonymized labels. Their real identities, stated here only: Code-render H = Hyperframes · Code-render R = Remotion · Pixel-diffusion L = LTX. The tables below use the real product names.

    A · Experiment protocol — versions, timeline, annotation ledger

    Timeline (2026). Briefs + fact checklists authored and frozen before Jul 6 → systems frozen Jul 6 → all generation Jul 6–12 → VBench-Long scoring Jul 13–14 → human annotation Jul 13–15 (first response Jul 13 09:41, last Jul 15 04:33) → statistics and human↔VBench alignment Jul 14–16. VBench scoring completed before any human-label comparison existed.

    Version pins — locally-run tools at their latest release as of the Jul 6 freeze; hosted platforms at the latest flagship model during the window. All 16:9, maximum quality settings, storyboard/planning features as shipped by each platform.

    SystemVersion · platform · settingsSource
    TapVidinternal build current at the access window—
    Fable 5Claude Code 2.1.202 · model claude-fable-5 (announced Jun 9, redeployed Jul 1) · defaultsannouncement · changelog
    Hyperframesv0.7.37 · Claude Code · Sonnet 5 (released Jun 30) · defaultsrelease · Sonnet 5
    Remotionv4.0.485 · Codex CLI 0.142.5 · GPT-5.5 (announced Apr 23; GPT-5.6 shipped Jul 9, after the freeze) · defaultsrelease · Codex · GPT-5.5
    VeoVeo 3.1 (announced 2025-10-15) · Google Flow, storyboard planning by Flow · Quality mode, 720pannouncement · modes doc
    LTXLTX-2.3 (released Mar 2026) · LTX Studio, storyboard planning by LTX Studio using GPT-Image-2 · Pro mode, 1080prelease · GPT-Image-2
    SeedanceSeedance 2.0 (announced Feb 12) · Jimeng — ByteDance’s official consumer platform — Agent Mode, story-film skill · VIP mode (top quality), 720pannouncement · Jimeng

    Prompt provenance. All 25 briefs and their fact checklists were authored with Claude Opus on Claude Code, human-reviewed, and frozen before generation began. If AI authorship biases the briefs toward any system’s idiom, the beneficiary would be Fable 5 — a competitor — not TapVid.

    The 6,760-annotation ledger. One row per annotator × scene × system × item: fact-check 2,326 (168 facts) + TeachQuiz comprehension 2,062 (149 questions) + T2VWorldBench four dimensions 1,384 (346 video×rater units × 4) + text accuracy 692 (346 units × grade + junk flag) + side-by-side preference 296 = 6,760 . Every count is recomputable from the raw export.

    VBench run. Official VBench-Long code ( vbench2_beta_long , commit e946a8d ) on 5× NVIDIA L4, one shard per use case, ~6 h wall; environment pinned and reproduced from scratch on a fresh instance; per-cell result JSONs and run logs retained. Input normalization and the >3.7M evaluation arithmetic: appendix E .

    B · Worked example — one brief through the whole pipeline

    The brief (How-To scene 1): *“The Paper That Pops — A No-Glue Origami Fidget Toy”* , target 5:00, 16:9. Like every brief, it is a fully-specified production document, identical for all seven systems: art direction with an exact hex palette, lighting and camera language, a nine-scene shot list with full voice-over text and per-scene motion notes, and a single master prompt — e.g. *“…the three color-coded papers (indigo 6×6 outer box, amber 5×5 inside button, violet 9×4 spring), the shared blintz fold that builds both boxes… the no-glue lock where flaps tuck under adjacent edges…”* The full information payload lives in the brief, so coverage failures are attributable to the system, not the prompt.

    Its required-fact checklist (authored with the brief, frozen before generation; station items rendered in English). A fact is credited only if it appears on screen verbatim and legible — annotators may replay frame-by-frame:

    • Outer box paper: 6×6 cm
    • Inner button paper: 5×5 cm
    • Spring paper: 9×4 cm
    • The no-glue promise: no glue, no tape, anywhere
    • The key fold, by name: blintz fold (corners to center)

    A TeachQuiz comprehension item (answered after viewing; the “video did not provide this” escape is scored incorrect): *“What is the key fold that both boxes share?”* — valley fold / blintz fold (corners to center) / squash fold / the video did not provide this information.

    How it scored. On this scene both of TapVid’s raters credited 5/5 facts and answered 12/12 comprehension questions correctly. The same items caught real misses elsewhere in the field on this scene — including the fold-name question — which is the discrimination the benchmark is built to measure.

    C · Statistical methods — estimators, seeds, and independent re-verification

    Estimators. Binary stations (coverage, delivery): Bayesian mixed logistic regression value ~ system + (1|scene) + (1|item) , TapVid as reference. Graded stations (text, dimensions): scene-cluster-robust OLS with annotator fixed effects. Delivery additionally via leave-one-out information gain. Three independent estimators — raw rates, information gain, mixed logit — produce the identical system ordering in both frames.

    Uncertainty. 95% CIs from scene-level bootstrap (resample the 25 scenes, B = 3000, fixed seed); significance = the 95% interval excludes zero; results reported as tiers where intervals overlap, never false-precision ranks. VBench: Skillings–Mack omnibus per dimension (method differences significant at p = 8e-18 to 1e-9 over 23 complete scene blocks), post-hoc pairwise Wilcoxon signed-rank with Holm correction. Human↔VBench alignment: Spearman rank correlation, stress-tested by 2000× bootstrap, leave-one-system-out, and permutation tests (all seeded).

    Reproducibility — independently re-verified. All analysis randomness is seeded, raw data (the full annotation export and per-scene VBench scores) ships with the analysis, and each metric has its own one-command reproduce script under pinned dependencies (Python 3.14.5 · numpy 2.5.1 · pandas 3.0.3 · scipy 1.18.0 · statsmodels 0.14.6 · matplotlib 3.11.0 · openpyxl 3.1.5). An independent re-run in a fresh environment regenerated all 55 committed result artifacts: every deterministic output byte-identical, all 15 figures pixel-identical, and the variational mixed-model columns matching to ≤6e-06 log-odds — with no significance verdict changing anywhere .

    D · Reliability, method, and limitations

    Rater robustness by design. The panel is fully crossed — every annotator rated every system — so rater strictness cannot bias a between-system comparison. Adding an annotator random effect moves contrasts by at most 0.16 (delivery) / 0.31 (coverage) log-odds with no significance flips.

    Inter-rater agreement per metric. Delivery: 73.6% raw agreement, Gwet AC1 0.59. Coverage: 67.2%, AC1 0.42 — the annotator effect is the dominant nuisance source (sd 0.98), so the absolute coverage level is rater-sensitive even though the comparison is not. Text: within-one-grade 91%, weighted κ 0.27 — the noisiest rating. WorldBench dimensions: within-one 79–85%, weighted κ 0.24–0.35. Aesthetic Preference: weighted κ 0.07 — aggregate win-rates over 296 comparisons are meaningful; individual judgments are not, which also attenuates the orthogonality correlations toward zero.

    Limitations. Delivery is relative to this seven-system field and not comparable across benchmark runs with a different field. Comprehension is administered open-book; the 0.86 median ceiling compresses system differences, and a closed-book run is the highest-leverage sharpening. 25 scenes (5 per use case) with 2 raters per video bound per-use-case resolution — the top tiers do not fully separate. Coverage credits a fact only when verbatim and legible. Mixed-logit intervals are variational (anti-conservative); the large significant contrasts are safe. The arena’s star topology supports no competitor-vs-competitor ranking and no absolute aesthetic score. T-scores are field-relative conveniences, not absolute grades. A facts-per-second bandwidth station is defined but was not computed in this round.

    E · VBench ↔ human alignment — method

    Instruments. VBench Frame Quality is the mean of the model’s Aesthetic Quality and Imaging Quality dimensions, computed quality-when-it-works (a failed generation is excluded, not scored 0). Human preference is the side-by-side win-rate over TapVid. The correlation is over the six competitor systems — TapVid is the star-topology anchor of every comparison and so is not a plotted point.

    Alignment (VBench Frame Quality ↔ human preference)Spearman ρ
    Overall+0.89
    Leave-one-system-out (worst case)+0.87
    Professional+0.94
    News+0.77
    Marketing+0.70
    How-To+0.48
    Explainer−0.33

    The overall sign survives 2000× bootstrap resampling and leave-one-system-out; n = 6 systems bounds formal significance. Explainer is the one segment where viewers reward motion over frame-polish, so the frame-quality metric tracks a different cue there. Every VBench composite was recomputed from the raw per-scene scores and matches the source report to the decimal.

    Run provenance and the >3.7M count. VBench-Long ( vbench2_beta_long at commit e946a8d ; VBench++, IEEE TPAMI 2025) was run on NVIDIA L4 GPUs on Jul 13–14, 2026 — before any human-label comparison existed. Inputs were normalized per scene so no system gains an encoding artifact: 720p cap (never upscaled), one uniform 24 fps frame grid, near-lossless re-encode. Six dimensions per video; four of them score *every frame* (24 model forwards/s each), motion smoothness runs 11.5/s and dynamic degree 7.5/s — 116 model forwards per second of footage. Across the 32,141 seconds of scored video: 3.73 million frame-level model evaluations . A further consistency re-check on length-trimmed builds (not counted above) verified that video length does not confound the comparison: native−trim deltas all < 0.0025.

    F · Composite scorecard — Informative Quality and the three axes

    Field-relative T-scores: 70 = the mean of these seven systems, every 10 points = 1 SD. Informative Quality = equal-weight mean of the Content Fidelity and Visual Quality axes. Aesthetic Preference is arena-derived and reported beside the informative axes, never inside them — it runs inverse to them. Overall Quality, (Content + Visual + Aesthetic Preference)/3, is the named information-dense scenario headline.

    All 25 scenes

    SystemInformative Quality95% CIContentVisualAesthetic Preference
    Fable 585.4[78.3, 92.1]82.688.252.3
    TapVid78.2[69.3, 86.6]77.179.261.7
    Remotion72.3[61.5, 82.2]76.668.169.2
    Hyperframes71.8[59.3, 83.0]73.370.268.1
    Veo64.8[51.7, 77.0]66.063.674.6
    Seedance61.3[46.8, 75.1]60.362.283.1
    LTX56.3[41.8, 70.8]54.158.581.0

    Information-dense scenes (Professional, News, How-To, Explainer)

    SystemOverall Quality (C+V+P)/3ContentVisualAesthetic PreferenceInformative Quality
    TapVid75.283.680.361.782.0
    Fable 575.082.888.154.285.5
    Seedance73.571.769.779.170.7
    LTX70.060.274.875.167.5
    Remotion69.777.565.965.771.7
    Hyperframes69.072.968.365.770.6
    Veo68.471.061.772.466.3

    Bootstrap CIs over the 25 scenes overlap widely: the tiers the data supports are Fable 5 · TapVid · Remotion/Hyperframes · the pixel trio. The Overall Quality TapVid–Fable 5 gap (75.2 vs 75.0) is within noise. The ranking is robust to method (PCA vs axis-mean), reliability weighting, and where Relevance is grouped. TapVid’s Aesthetic Preference is 0.5-anchored by the arena’s star topology; its per-use-case variation lives in the competitors.

    G · Factual Coverage — both frames

    All 25 scenes — share of required facts on screen, verbatim and legible; 95% bootstrap CI; mixed-logit contrast.

    SystemCoverage95% CIΔ vs TapVid95% CIResult
    Fable 50.790[0.717, 0.860]+0.094[−0.183, +0.370]TIE
    TapVid0.773[0.699, 0.843]—
    Remotion0.759[0.668, 0.845]−0.166[−0.417, +0.085]TIE
    Hyperframes0.745[0.642, 0.842]−0.199[−0.448, +0.050]TIE
    Seedance0.611[0.506, 0.710]−0.826[−1.052, −0.601]WIN
    Veo0.596[0.494, 0.697]−0.867[−1.091, −0.642]WIN
    LTX0.541[0.433, 0.658]−1.179[−1.400, −0.958]WIN

    Information-dense scenes

    SystemCoverage95% CIΔ vs TapVid95% CIResult
    TapVid0.836[0.773, 0.890]—
    Fable 50.792[0.710, 0.863]−0.291[−0.601, +0.020]TIE
    Remotion0.765[0.669, 0.857]−0.520[−0.801, −0.239]WIN
    Hyperframes0.727[0.607, 0.835]−0.681[−0.952, −0.409]WIN
    Seedance0.677[0.578, 0.770]−0.922[−1.182, −0.662]WIN
    Veo0.601[0.497, 0.714]−1.213[−1.464, −0.963]WIN
    LTX0.545[0.418, 0.671]−1.553[−1.798, −1.307]WIN

    Coverage ↔ delivery: Pearson r = 0.70 across 173 videos; pooled OLS gives P(comprehended | on screen) = 0.96 and a 0.36 off-screen floor; a mixed logit agrees (0.945 / 0.30). Coverage’s absolute level is rater-sensitive (see D); the between-system comparison is not.

    H · TeachQuiz Information Delivery — both frames

    All 25 scenes — raw comprehension correct-rate, mixed-logit predicted P(correct) at a typical scene/question, and contrast.

    SystemRaw ratePredicted PΔ vs TapVid95% CIResult
    Fable 50.8210.868+0.243[−0.091, +0.576]TIE
    Remotion0.8120.857+0.147[−0.167, +0.461]TIE
    TapVid0.7950.838—
    Hyperframes0.7890.834−0.028[−0.330, +0.274]TIE
    Veo0.7850.830−0.052[−0.353, +0.249]TIE
    LTX0.7150.756−0.509[−0.786, −0.233]WIN
    Seedance0.6910.730−0.647[−0.918, −0.377]WIN

    Information-dense scenes

    SystemRaw rateΔ vs TapVid95% CIResult
    Fable 50.846+0.056[−0.342, +0.453]TIE
    TapVid0.845—
    Veo0.840+0.005[−0.368, +0.378]TIE
    Remotion0.832−0.067[−0.433, +0.299]TIE
    Hyperframes0.790−0.393[−0.731, −0.054]WIN
    Seedance0.786−0.423[−0.759, −0.087]WIN
    LTX0.752−0.650[−0.971, −0.329]WIN

    Delivery is relative to this seven-system field (leave-one-out gain / field-baselined logit); administration is open-book, which compresses differences toward the 0.86 median ceiling. Three estimators (raw, gain D, mixed logit) agree on the ordering in both frames.

    I · T2VTextBench Text Accuracy and T2VWorldBench dimensions

    T2VTextBench Text Accuracy — four-tier grade on [0,1]; cluster-robust contrast vs TapVid. * = 95% CI excludes zero.

    SystemAll 2595% CIInfo-dense95% CIΔ all-25
    Fable 50.783[0.714, 0.850]0.750[0.681, 0.825]+0.062
    TapVid0.740[0.660, 0.820]0.769[0.675, 0.856]—
    Hyperframes0.700[0.615, 0.785]0.706[0.613, 0.794]−0.009
    Remotion0.695[0.610, 0.785]0.688[0.594, 0.781]−0.050
    Seedance0.640[0.535, 0.740]0.700[0.594, 0.806]−0.087
    Veo0.630[0.530, 0.730]0.650[0.531, 0.762]−0.090
    LTX0.580[0.500, 0.660]0.619[0.537, 0.694]−0.145*

    T2VWorldBench, all 25 scenes — mean grade per dimension and the paper composite. * = significant vs TapVid (cluster-robust OLS).

    SystemQualityRealismRelevanceConsistencyComposite
    Fable 50.900*0.9090.8910.9090.902*
    TapVid0.8600.8800.8440.8720.864
    Remotion0.8000.8600.8600.8240.836
    Hyperframes0.8120.8520.8160.8440.831
    Veo0.8000.8040.8080.8240.809
    Seedance0.8080.8080.772*0.796*0.796*
    LTX0.7920.784*0.688*0.792*0.764*

    The composite is internally valid (Cronbach α = 0.91) but folds one content-facing axis: Relevance correlates with Delivery/Coverage (+0.77) more than with its sibling visual dimensions (0.61–0.66). Info-dense frame: TapVid stays 2nd on every dimension; only the Quality and Composite gaps to Fable 5 remain significant.

    J · Aesthetic Preference (arena-style) — both frames, both readings

    Star topology: every comparison is TapVid vs one competitor on the same brief (296 comparisons). Win-rate counts ties as half; non-loss counts ties for TapVid; both are reported. Verdicts from binomial mixed logits vs the 0.5 no-preference line.

    All 25 scenes

    vsWin-rate95% CINon-loss95% CIVerdict
    Fable 50.587[0.482, 0.683]0.696[0.577, 0.808]TapVid does-not-lose*
    Hyperframes0.440[0.347, 0.531]0.560[0.458, 0.667]toss-up
    Remotion0.430[0.338, 0.519]0.540[0.433, 0.643]toss-up
    Veo0.380[0.297, 0.464]0.480[0.389, 0.571]toss-up
    LTX0.320[0.217, 0.417]0.380[0.267, 0.500]TapVid loses*
    Seedance0.300[0.206, 0.393]0.400[0.286, 0.500]TapVid loses*
    Overall0.407[0.35, 0.47]0.507[0.44, 0.57]less preferred than the field*

    Information-dense scenes — overall win-rate 0.432 [0.37, 0.50], non-loss 0.530: parity. The significant non-loss deficits vanish; on strict decisive comparisons only Seedance remains significantly preferred.

    vsWin-rate95% CINon-loss95% CI
    Fable 50.569[0.450, 0.692]0.667[0.542, 0.800]
    Remotion0.463[0.354, 0.568]0.575[0.458, 0.692]
    Hyperframes0.463[0.361, 0.563]0.600[0.500, 0.714]
    Veo0.400[0.304, 0.500]0.500[0.385, 0.607]
    LTX0.375[0.250, 0.500]0.425[0.292, 0.545]
    Seedance0.338[0.225, 0.450]0.425[0.286, 0.567]

    Preference is empirically orthogonal to the informative metrics: correlating TapVid’s per-video advantage on each metric with the arena outcome gives −0.22 to −0.01 (n = 148 cells) — all negative or near zero. Non-loss is the favourable reading (20% of comparisons are ties); both are shown for that reason.

    K · Per-use-case standings — where TapVid wins and loses

    Significant wins/losses vs TapVid per use case (95% interval excludes zero). Cells list competitors; “—” = none.

    Use caseCoverage: TapVid beatsloses toDelivery: TapVid beatsloses to
    ProfessionalFable 5, Hyperframes, Veo, LTX—Fable 5, LTXSeedance
    NewsRemotion, Veo, LTX, Seedance—Remotion, LTX, Seedance—
    How-ToRemotion, Hyperframes, Veo, LTX—Veo—
    ExplainerLTX, SeedanceFable 5SeedanceFable 5
    MarketingSeedanceFable 5, Remotion, HyperframesSeedanceFable 5, Remotion, Hyperframes

    No system wins everywhere: Seedance leads Professional delivery yet collapses on Marketing (P 0.30); Veo ties the News delivery lead yet is worst-in-field on How-To. TapVid’s arena win-rate by use case: Professional 0.450, News 0.433, Explainer 0.429, How-To 0.417, Marketing 0.308. TapVid’s composite Informative Quality by use case: News 87, How-To 89, Explainer 81, Professional 71, Marketing 63.

    10

    Try TapVid

    Turn supplied assets and approved copy into a reviewable explainer-video workflow. Explore TapVid.

    11

    References and citation

    Supporting evaluation methods include T2VTextBench, T2VWorldBench, and the Text-to-Video Human Evaluation protocol.

    Recommended citation (APA): Law, D. (2026, September 6). AI video benchmark: Accuracy vs aesthetics (Version 1.0). TapVid Research. https://tapvid.ai/blog/information-dense-ai-video-benchmark

    @techreport{law2026aiVideoBenchmark,
      author = {Derek Law},
      title = {AI Video Benchmark: Accuracy vs Aesthetics},
      institution = {TapVid Research},
      year = {2026},
      month = {September},
      url = {https://tapvid.ai/blog/information-dense-ai-video-benchmark},
      note = {Version 1.0; July 2026 benchmark run}
    }

    12

    About the author

    Derek Law
    Derek Law

    Derek Law is CTO and co-founder of TapVid. His work is guided by a simple thesis: the constraint on AI is the interface, not intelligence. Previously, he worked on Alexa AI and GenUI for Amazon AGI.

    How this article was verified

    Basis: TapVid's editorial publishing record for this articleEvidence: Author attribution, publication history, and source links where external claims require them

    Article versionSeptember 7, 2026

    About the authorDerek Law

    CTO & Co-founder, TapVid

    Derek Law is CTO and co-founder of TapVid. His work is guided by a simple thesis: the constraint on AI is the interface, not intelligence. Previously, he worked on Alexa AI and GenUI for Amazon AGI.

    View all 1 articles →

    Derek Law invites you to join the conversation with fellow video creators on Discord.

    Join Derek on Discord →
    Turn approved assets and copy into a reviewable explainer video

    Use the materials you already have

    From yourfilesfilesto a ready-to-publish video

    WEB→ VIDEOPPT→ VIDEOPDF→ VIDEOASSETS→ VIDEOAUDIO→ VIDEOVIDEO→ VIDEOTALKING HEAD→ VIDEOWEB→ VIDEOPPT→ VIDEOPDF→ VIDEOASSETS→ VIDEOAUDIO→ VIDEOVIDEO→ VIDEOTALKING HEAD→ VIDEO

    Keep reading

    Related stories

    Video Completion Rate 2026: Seven Definitions Audited research cover
    Research·11 min read

    Video Completion Rate 2026: Seven Definitions Audited

    Seven first-party definitions show how plays, starts, impressions, viewers, and milestone counts change video completion rate.

    Sep 4, 2026

    AI video workflow benchmark based on 6,464 matched TapVid users
    Research·11 min read

    AI Video Workflow Benchmark: 6,464 Users Analyzed

    A matched cohort of 6,464 TapVid users shows why brief, script, valid video, viewing, export, and sharing need separate workflow metrics.

    Sep 2, 2026

    Make an animated video fast without losing quality, with TapVid
    How-to·10 min read

    How to Make an Animated Video Fast Without Losing Quality

    A fast, quality-safe method for teams and creators who need to make an animated video under real deadlines.

    Apr 16, 2026

    Ready to create your first video?

    Join thousands of product teams using AI to create professional videos in minutes.

    Your first video in under 5 minutes →Book a demo →
    Tapvid

    TapVid turns the materials your business already has into an accurate video that explains the job clearly and is ready to publish.

    TikTokInstagramXDiscordYouTube

    TapVid

    Features

    AI Explainer Video GeneratorAI Motion Graphics GeneratorAI Product Demo Video GeneratorProduct Demo Video MakerExplainer Video TemplatesVideo Production Plan TemplateVideo Creative Brief TemplateCorporate Video TemplateVideo Sales Letter TemplateVideo Production Proposal TemplatePromo Video TemplateVideo Production TemplateAI Product Video GeneratorAI B-Roll GeneratorTalking Head EditingClone VideoPrompt to VideoText to Video AIText to Motion GraphicsAnimated Video MakerAnimated Explainer Video MakerKinetic Typography GeneratorAnimated Chart MakerAnimated Collage MakerFree AI Video Generator

    Convert to Video

    Screenshot to VideoImage to VideoAssets to VideoAudio to VideoVideo to Video AIPDF to VideoPPT to VideoArticle to VideoBlog to VideoURL to VideoScript to VideoGoogle Slides to VideoWord to VideoSOP to Video

    Use Cases

    AI Study Video MakerSaaS Explainer VideoAI Video AutomationSaaS Video ProductionIndustrial Video ProductionProduct Launch Video MakerAI Ad Video GeneratorDocumentary Video MakerAnimated Social Media Video MakerInfographic Video MakerPodcast to VideoWhiteboard Animation MakerWhiteboard Explainer VideoEcommerce Video AdsStartup Explainer VideoCase Study VideoEducational VideoTutorial VideoCustomer OnboardingHelp Center VideoAPI Docs Video

    Solutions

    Explainer VideoProduct Demo VideoMeeting Recap VideoWebinar ClipsMarketing VideoFeature AnnouncementCompetitive ComparisonNewsletter VideoLanding Page VideoInvestor Pitch Video

    Featured Guides

    Video Prompt LibraryGemini Omni 1.1 Flash Prompt LibraryMiniMax H3 Prompt LibraryBest Faceless YouTube NichesCollage Animation Guide

    Company

    All FeaturesAboutBlogPricingGet in Touch

    © 2026 TapVid. All rights reserved.

    Privacy
    Terms of Service