eachlabs-voice-audio

Convert text to speech and transcribe audio with speaker diarization.

28|5|Updated Feb 9, 2026
One-click install
npx skills add https://github.com/eachlabs/skills --skill eachlabs-voice-audio
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: eachlabs-voice-audio
Source: https://github.com/eachlabs/skills/tree/main/skills/eachlabs-voice-audio
Command: npx skills add https://github.com/eachlabs/skills --skill eachlabs-voice-audio

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a comprehensive suite of tools for voice and audio manipulation, including text-to-speech, speech-to-text, and voice conversion, simplifying complex audio tasks.

Core Features & Use Cases

  • Text-to-Speech (TTS): Generate natural-sounding speech from text using various models like ElevenLabs and Kling.
  • Speech-to-Text (STT): Transcribe audio files with options for diarization (speaker identification) and timestamps using models like Whisper and Wizper.
  • Voice Conversion & Cloning: Transform voices or clone them using models like RVC and ElevenLabs.
  • Audio Utilities: Merge audio with video and perform general audio conversions.
  • Use Case: You need to add a voiceover to a video in multiple languages, transcribe a meeting with speaker labels, or create a custom AI voice for your brand.

Quick Start

Use the eachlabs-voice-audio skill to convert the text 'Hello, world!' into speech using the elevenlabs-text-to-speech model.

Frequently Asked Questions about eachlabs-voice-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio with speaker diarization and timestamps?▼

To transcribe audio with speaker labels, speech-to-text models like Whisper and Wizper process audio files to generate text output that includes speaker diarization and timestamps.

Can I generate natural speech from text using ElevenLabs?▼

Text-to-speech generation using ElevenLabs converts written text into natural-sounding speech, supporting multiple languages for voiceovers and diverse audio generation needs.

What's the best way to convert or clone a custom voice?▼

Voice conversion and voice cloning transform or replicate voices using models like RVC and ElevenLabs, allowing you to create a custom AI voice for your brand or project.

Do I need an API key to run text-to-speech and voice conversion predictions?▼

An API key is required for authentication to run text-to-speech and voice conversion predictions, integrating directly with the EachLabs API for execution.

How do I merge generated audio with an existing video file?▼

Audio utilities within the voice and audio processing suite allow you to merge generated audio tracks with existing video files and perform general audio format conversions.