What problem does it solve? Creating talking-head or presenter videos normally requires cameras, actors, and studio time. This Skill guides the production of realistic lip-synced AI avatar videos from just one portrait image and one audio file, using ByteDance's OmniHuman1 framework via platforms like SousakuAI, Fal.ai, and BytePlus. ## Core Features & Use Cases - Image + Audio to Video Workflow: Step-by-step guidance for preparing portrait images (real photos, anime characters, illustrations) and audio (speech, singing, dialogue) for generation. - Prompt Engineering for Avatars: Templates for camera work, gestures, expressions, and atmosphere to control the generated video's look and feel. - VSL Production Pipeline: End-to-end flow covering script writing, voice recording (or AI voices like VOICEVOX and ElevenLabs), avatar image generation, OmniHuman1 rendering, and post-editing in Canva or CapCut. - Use Case: A marketer needs a 7-minute video sales letter but has no presenter. They generate a professional avatar image, record or synthesize the narration, and produce a lip-synced presenter video with captions and a CTA screen. ## Quick Start Ask the assistant to walk you through creating an OmniHuman1 avatar video from your portrait image and narration audio, including the recommended SousakuAI settings and prompt.