What problem does it solve? Creating finished audio—narration, multi-voice dialogue, sound effects, or voice clones—normally requires stitching together separate TTS, cloning, and sound-design tools. This Skill turns a single text script or reference image into a completed audio file through one async API workflow, with billing, polling, and file download handled for you. ## Core Features & Use Cases - End-to-End Audio Generation: Submit a text script (up to 1400 characters) to the ListenHub-Voice-1.0 model and receive a finished audio track, including sound effects described inline in the text. - Multi-Voice Dialogue & Voice Cloning: Assign lines to 2–3 voices with @音频N prefixes, clone a voice from a public reference audio URL, or use a registered ListenHub speaker or official voice_type. - Image-to-Audio: Convert a reference image (URL or Base64) into an audio track, mutually exclusive with voice selection. - Use Case: Ask for a 15-second rain-and-thunder sound effect, or a two-person podcast dialogue cloned from two reference clips, and the Skill submits the task, polls until success, and saves the MP3 to your working directory. ## Quick Start Ask the assistant to generate an audio clip from your script using listenhub-voice, for example: generate a 15-second sound effect of steady rain with distant thunder.