In plain words: A new collection of 42 million hand-drawn cartoon keyframes, each labeled with descriptions and tags, gives AI models plenty of cartoon data to learn from instead of natural video. Fine-tuning video models on it made them much better at understanding and generating cartoons.
Abstract
Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural video by scaling large models with extensive datasets, they are not as effective for cartoons. Through our empirical experiments, we argue that this ineffectiveness stems from a notable bias in hand-drawn cartoons that diverges from the distribution of natural videos. Can we harness the success of the scaling paradigm to benefit cartoon research? Unfortunately, until now, there has not been a sizable cartoon dataset available for exploration. In this research, we propose the Sakuga-42M Dataset, the first large-scale cartoon animation dataset. Sakuga-42M comprises 42 million keyframes covering various artistic styles, regions, and years, with comprehensive semantic annotations including video-text description pairs, anime tags, content taxonomies, etc. We pioneer the benefits of such a large-scale cartoon dataset on comprehension and generation tasks by finetuning contemporary foundation models like Video CLIP, Video Mamba, and SVD, achieving outstanding performance on cartoon-related tasks. Our motivation is to introduce large-scaling to cartoon research and foster generalization and robustness in future cartoon applications. Dataset, Code, and Pretrained Models will be publicly available.
Zhenglin Pan
arXiv:2405.07425 · cs.CV · submitted May 13, 2024
abstract · pdf · arXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission
A large chunk of an anime's budget goes to in-betweening. Essentially human interpolation of graphics between two key frames (Its usually like 6-12 inbetweens per key frame). People hate this job, and it is highly unproductive, so generally outsourced by Japan to other countries.
Western animation decided to abandon it altogether, and move first to flash, then to 3d animation. But in retrospect that was a mistake, as it lost so much of the creative flexibility of 2d animation. Anime today is substantially bigger than western animation as a result. Crunchyroll has 13 million subscribers.
AI will solve the problems with 2d animation. Something like SORA fine-tuned on anime-keyframe data like this. Can probably easily solve in-betweening. Then the 2d animation workflow will dominate 3d. Its so much easier to just draw a beginning and end key-frame, then have the AI fill it in. Than to model and rig and render a entire scene.