AnimationBench: What It Establishes and Where Its Scope Stops
October 5, 2026 3 min read
What the paper is
AnimationBench, published on arXiv under the title "Are Video Models Good at Character-Centric Animation?", is a benchmark for evaluating image-to-video generation models on animation-style content. The paper positions it as the first systematic benchmark built specifically for animation rather than adapted from realism-oriented video evaluation. It grounds its criteria in two anchors: IP Preservation — whether a generated character stays consistent in appearance, behavior, and personality over time — and the Twelve Basic Principles of Animation from Johnston and Thomas (1981), operationalized into measurable dimensions such as squash and stretch, anticipation, slow in and slow out, arcs, secondary action, appeal, exaggeration, and dynamic degree. A third tier adds broader generative video qualities: semantic consistency, motion rationality, and camera motion consistency.
How it works
The design has two modes. A close-set protocol uses standardized prompts and curated character assets so different models can be compared reproducibly. An open-set mode takes any animation video and prompt, runs a diagnostic pass that identifies specific failure modes, and feeds findings into a prompt refiner — turning the benchmark into a tool for iterating on a particular generation rather than just ranking models. Scoring leans on visual-language models acting as structured judges: instead of asking whether a video is good, the pipeline asks targeted questions grounded in observable detail, which the authors argue keeps scores tied to concrete observations rather than subjective impressions. This is how the benchmark claims to cover stylized motion and character-centric consistency that realism-focused benchmarks tend to miss.
What it establishes
Within its own scope, the paper reports that the benchmark aligns well with human judgment and exposes animation-specific quality differences that realism-oriented evaluation overlooks, making it more discriminative for state-of-the-art image-to-video models. The practical takeaway is narrower than the headline question: current models handle general video quality reasonably but struggle with expressive, principle-driven character motion and long-term identity preservation. That gap between "consistent video" and "consistent character" is the finding the paper exists to make visible.
Where the scope stops
Three limits matter for anyone applying these results. First, the benchmark measures visual animation quality; it does not assess synchronization between sound, speech, and visual motion, so it cannot speak to multimodal generation where audio alignment is part of the task. Second, the character-preservation metric is a measurement of consistency, not a determination of intellectual-property rights — a high score says nothing about whether a character depiction infringes someone else's IP. Third, the evaluation is VLM-mediated, so its conclusions are as strong as the judge models behind them, and its "first systematic benchmark" claim is the authors' own positioning, not an independently verified one. The checked version is an author-submitted arXiv paper.
When the method applies
Use AnimationBench when your question is specifically about animation-quality image-to-video generation: does the model keep a character on-model, does its motion read as animated rather than merely smooth, and where exactly does a given output fall short? If you are evaluating photorealistic video, general text-to-video composition, or anything involving audio, the benchmark's dimensions do not answer those questions, and its absence of coverage there is a scope boundary rather than evidence of failure.
Sources
- AnimationBench: Are Video Models Good at Character-Centric Animation? — arXiv (author-submitted research)