audio-gen

Generate multilingual speech, voice clones, and sound effects via ElevenLabs TTS.

267|63|Updated Feb 15, 2026
One-click install
npx skills add https://github.com/modu-ai/cowork-plugins --skill audio-gen
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: audio-gen
Source: https://github.com/modu-ai/cowork-plugins/tree/main/moai-media/skills/audio-gen
Command: npx skills add https://github.com/modu-ai/cowork-plugins --skill audio-gen

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires elevenlabs, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables seamless AI-driven audio production by generating realistic voices, dubbing, and sound effects to streamline multimedia content creation.

Core Features & Use Cases

  • Text-to-Speech and Voice Cloning: Produce multilingual narration and replicate voices from short samples for branding or personalization.
  • Multilingual Dubbing: Automatically translate and synchronize audio tracks for videos in multiple languages, saving time on manual dubbing.
  • Sound Effect Generation: Create custom sound effects for movies, games, or podcasts based on descriptive prompts.
  • Use Case: You can generate a professional narration in Korean, clone a voice from a 1-minute sample, and dub a video into English or Japanese with LipSync adjustments.

Quick Start

Input your text prompt requesting an audio narration in your preferred language, select a voice preset, and then generate the audio file for immediate use.

Frequently Asked Questions about audio-gen

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate multilingual speech for video dubbing?▼

Voice cloning replicates a specific voice from a short one-minute audio sample for branding or personalization. You provide the sample, and the system generates new narration matching the original voice characteristics.

Can I create custom sound effects from text prompts?▼

You can create custom sound effects by inputting descriptive text prompts. The system generates tailored audio assets for movies, games, or podcasts based directly on your provided descriptions.

Do I need Python libraries to use ElevenLabs text-to-speech?▼

You need Python libraries for audio processing and voice synthesis to use this ElevenLabs text-to-speech integration. These dependencies handle the underlying audio generation and file output operations.

What is the best way to clone a voice from a short audio sample?▼

The best way to clone a voice is providing a clear one-minute audio sample to the system. It replicates the voice characteristics to produce new narration for branding or personalization.