omnihuman1-video

Generates lip-synced AI avatar videos from a single image and audio file using OmniHuman1.

Updated Mar 23, 2026
One-click install
npx skills add https://github.com/sakamotomomotaro0809-netizen/tateyomi --skill omnihuman1-video-sakamotomomotaro0809-netizen
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: omnihuman1-video
Source: https://github.com/sakamotomomotaro0809-netizen/tateyomi/tree/main/taisun_agent/.claude/skills/omnihuman1-video
Command: npx skills add https://github.com/sakamotomomotaro0809-netizen/tateyomi --skill omnihuman1-video-sakamotomomotaro0809-netizen

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve? Creating talking-head or presenter videos normally requires cameras, actors, and studio time. This Skill guides the production of realistic lip-synced AI avatar videos from just one portrait image and one audio file, using ByteDance's OmniHuman1 framework via platforms like SousakuAI, Fal.ai, and BytePlus. ## Core Features & Use Cases - Image + Audio to Video Workflow: Step-by-step guidance for preparing portrait images (real photos, anime characters, illustrations) and audio (speech, singing, dialogue) for generation. - Prompt Engineering for Avatars: Templates for camera work, gestures, expressions, and atmosphere to control the generated video's look and feel. - VSL Production Pipeline: End-to-end flow covering script writing, voice recording (or AI voices like VOICEVOX and ElevenLabs), avatar image generation, OmniHuman1 rendering, and post-editing in Canva or CapCut. - Use Case: A marketer needs a 7-minute video sales letter but has no presenter. They generate a professional avatar image, record or synthesize the narration, and produce a lip-synced presenter video with captions and a CTA screen. ## Quick Start Ask the assistant to walk you through creating an OmniHuman1 avatar video from your portrait image and narration audio, including the recommended SousakuAI settings and prompt.

Frequently Asked Questions about omnihuman1-video

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create a lip-sync AI avatar video with OmniHuman1?▼

Upload one portrait image and one audio file to a platform hosting OmniHuman1, such as SousakuAI, optionally add a prompt describing camera work and gestures, then start generation. Processing typically takes 1-5 minutes and outputs an MP4 video.

What image and audio formats does OmniHuman1 support?▼

Images should be PNG or JPG at 1024x1024 pixels or higher, with a clear front-facing subject. Audio should be MP3 or WAV at 44.1kHz and 16-bit, with a recommended length of 30 seconds to 7 minutes.

Which platforms can I use to access OmniHuman1?▼

OmniHuman1 is available through SousakuAI, which offers a Japanese-friendly UI and OmniHuman 1.5 support, Fal.ai for API-based developer access, and BytePlus for enterprise-scale official service.

Why does my OmniHuman1 video have unnatural mouth movement?▼

Unnatural lip-sync usually comes from low-quality audio. Remove background noise, ensure clear pronunciation, and keep a moderate speaking pace of roughly 150-180 characters per minute before regenerating.

Can I use AI-generated voices with OmniHuman1 videos?▼

Yes, AI voice tools like VOICEVOX, ElevenLabs, and CoeFont can produce the narration audio. VOICEVOX is free and permits commercial use, while ElevenLabs offers richer emotional expression at a cost.

What are the ethical requirements for AI avatar videos?▼

You must disclose that the video is AI-generated, never use real people's likenesses without permission, and avoid deepfake misuse. Use original or licensed images and verify commercial licensing terms before publishing.