What problem does it solve? Turning a spoken line or topic into a finished vertical video normally requires manual motion design, timeline editing, and repeated trial-and-error with video models. This Skill compresses that into a gated pipeline: design a visual metaphor, generate a still frame, then animate it with first/last-frame interpolation, with QA artifacts and cost tracking at every stage. ## Core Features & Use Cases - Single or batch 5-second B-roll: Convert each spoken line into one visual metaphor (max 4 object groups, semantic background color), generate a halftone paper-collage still via Gemini image models, then animate assemble-from-empty motion via Gemini Omni or fal.ai models (Kling, Seedance). - Full 45-60 second explainer videos: Build a beat map (8-10 beats with narrative roles from 14 arc templates), generate per-beat TTS narration, role-based tempo adjustment, Pillow-rendered captions and headline cards, then assemble everything with a single ffmpeg pass driven by beats.json. - Cost and quality discipline: Environment self-check, resume-only-failed-jobs dispatch, automatic cost ledger, objective QA scores (first-frame purity, end-frame similarity), and deterministic HyperFrames routing for rigid-body motions to cut generation cost. - Use Case: Give it three voiceover lines about workflow automation and receive three silent 9:16 MP4 clips with contact sheets, first-frame verification, end-frame comparisons, and QA documents—ready to drop into a video edit. ## Quick Start Ask the agent to turn this spoken line into an editorial paper-collage B-roll clip and confirm the metaphor before any video generation starts.