What problem does it solve? Creating cinematic short-form video with synchronized lip-synced audio normally requires stitching together separate generation, TTS, and editing tools. This Skill routes video generation requests to ByteDance Seedance 2.0 Pro on the RunComfy Model API, handling multi-modal references (images, videos, audio) in a single CLI call. ## Core Features & Use Cases - Multi-modal video generation: Combine up to 9 image references, 3 video clips, and 3 audio references with a text prompt to produce 4–15 second videos at 480p or 720p. - Native lip-synced audio: Generate in-pass synchronized speech, SFX, and music with generate_audio, ideal for spokesperson ads and dialogue content. - Model routing guidance: Built-in decision table for when to use Seedance 2.0 Pro versus HappyHorse 1.0, Wan 2.7, Kling, or LTX 2. - Use Case: A marketer needs a 9:16 lip-synced ad where a specific barista explains today's special. Provide the headshot as image_url, describe the scene and tone in the prompt, and receive an 8-second vertical video with natural lip-sync. ## Quick Start Ask the agent to generate a short video with Seedance 2.0 Pro, for example: "Use Seedance to make a 5-second video of a barista on a cafe terrace at golden hour explaining today's special with natural lip-sync."