ai-avatar-video

Create avatar videos from audio or scripts via RunComfy model routing.

31|9|Updated Apr 30, 2026
One-click install
npx skills add https://github.com/agentspace-so/runcomfy-agent-skills --skill ai-avatar-video-agentspace-so
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: ai-avatar-video
Source: https://github.com/agentspace-so/runcomfy-agent-skills/tree/main/ai-avatar-video
Command: npx skills add https://github.com/agentspace-so/runcomfy-agent-skills --skill ai-avatar-video-agentspace-so

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It solves the problem of quickly creating realistic avatar and talking-head videos from a voiceover or a script, without manually juggling multiple tools and model-specific inputs.

Core Features & Use Cases

  • Model routing by intent: chooses the best RunComfy avatar route (OmniHuman, Wan 2-7 with audio_url, HappyHorse, Seedance v2 Pro, or Wan 2-2 Animate) based on whether you have an audio file or only text and what style you want.
  • Audio-driven lip-sync options: supports portrait+MP3 lip-sync (OmniHuman), prompt+audio_url scene generation (Wan 2-7), and script-driven in-pass speech (HappyHorse).
  • Cinematic, multi-modal generation: composes reference subject visuals and reference audio for more cinematic outcomes (Seedance v2 Pro).

Quick Start

Ask your AI to generate an audio-driven avatar video by providing a portrait image URL and a voiceover audio URL, then run the skill with those inputs so it selects the correct RunComfy route automatically.

Frequently Asked Questions about ai-avatar-video

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create a talking-head video from a portrait and an audio file?▼

To create a talking-head video, provide a portrait image URL and a voiceover audio URL, and the skill routes your inputs to the appropriate RunComfy avatar model for portrait lip-sync. It generates a video matching the audio track to the portrait.

Can I generate an avatar video from just a text script without a pre-recorded MP3?▼

Yes, you can generate an avatar video from a text script using the script-to-video speech-insertion route. This synthesizes speech from your script and drives the character animation without requiring a separate MP3 file.

What's the best way to achieve cinematic avatar generation with reference visuals?▼

For cinematic avatar generation, use the Seedance v2 Pro route to compose reference subject visuals and reference audio. This produces more cinematic outcomes by combining multi-modal reference inputs.

Do I need the RunComfy CLI to use audio-driven lip sync models?▼

Yes, you need to use the RunComfy CLI with the correct model endpoints and JSON input schema for image_url, audio_url, and prompt fields. The CLI is required to route your request to the correct avatar model.

How does the model routing decide which RunComfy route to use for my video?▼

Model routing selects the RunComfy avatar route by checking whether you have an audio file or only text, and your desired style. It automatically chooses between OmniHuman, Wan 2-7, HappyHorse, Seedance v2 Pro, or Wan 2-2 Animate.

Can I generate stylized character animation from an image using prompt and audio_url?▼

Yes, you can generate stylized character animation by providing a prompt and an audio_url to the Wan 2-7 route. This creates a scene driven by both your text prompt and the provided audio track.