What problem does it solve?
This Skill solves the problem of producing high-quality Kling 3.0 multi-shot videos with consistent character identity, synchronized audio, and correct tier/mode routing without manually orchestrating multiple API calls.
Core Features & Use Cases
- Six Kling 3.0 endpoints (tier × mode): text-to-video and image-to-video across Standard (1080p), Pro (1080p), and 4K (3840x2160) rendering tiers.
- Multi-segment prompting for multi-shot scenes: supports single-call, numbered shot sequences to preserve identity across shots.
- Optional synchronized audio generation: enables
generate_audio for Standard and Pro with higher cost and flat-rate audio behavior on the 4K tier.
- i2v image-based animation: animates a subject from a publicly fetchable HTTPS image URL, with optional
tail_image_url for controlled endings.
Quick Start
Use the Kling 3.0 skill to generate a 10-second 16:9 text-to-video by describing your cinematic multi-shot prompt and choosing the appropriate tier (Standard, Pro, or 4K).