mofa-podcast

Convert topics or scripts into multi-speaker podcast MP3s with TTS voices.

11|12|Updated Feb 28, 2026
One-click install
npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-podcast
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: mofa-podcast
Source: https://github.com/mofa-org/mofa-skills/tree/main/mofa-podcast
Command: npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-podcast

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

Generating professional multi-speaker podcasts from topic briefs or texts without expensive recording setups or editing.

Core Features & Use Cases

  • Script-driven podcast generation from topic or provided text with emotion cues and music guidelines.
  • Supports built-in voices plus custom/cloned voices via mofa-fm, with per-segment timeline assembly and BGM.
  • Outputs ready-to-publish MP3s and segment audio for post-production workflows, scalable for episodes and shows.

Quick Start

Provide a topic or markdown script to generate a complete multi-speaker podcast using built-in or cloned voices.

Frequently Asked Questions about mofa-podcast

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a multi-speaker podcast from a text script?▼

You generate a multi-speaker podcast by providing a topic or an approved markdown script, which the system transforms into a scripted audio show using TTS voices, emotion cues, and music guidelines.

Can I use cloned voices for podcast audio assembly?▼

Yes, you can use cloned voices for podcast audio assembly. The system supports built-in voices and custom voice cloning via mofa-fm to assign distinct voices to 1-5 speakers per segment.

Do I need FFmpeg to output podcast MP3 files?▼

You do not strictly need FFmpeg to output podcast audio. The system outputs a final MP3 file when FFmpeg is available, but it defaults to generating WAV files if FFmpeg is missing from your environment.

How many speakers can I include in a single TTS podcast generation?▼

You can include between 1 and 5 speakers in a single TTS podcast generation. The system supports multi-speaker configurations suitable for short-form podcasts, interviews, talk shows, and educational narrations.

What is the best way to add background music and emotion cues to TTS audio?▼

The best way to add background music and emotion cues is by using an approved markdown script format. The system processes these script guidelines to drive per-segment timeline assembly and BGM integration during audio generation.

What audio formats are outputted for post-production podcast workflows?▼

The audio formats outputted for post-production workflows are MP3 and WAV. The system generates ready-to-publish MP3s by default and provides individual segment audio files to allow further editing.