fal-ai-media

Generate images, videos, and audio via fal.ai MCP tools.

Updated Mar 20, 2026
One-click install
npx skills add https://github.com/KanakMalpani/General-Private-Skills --skill fal-ai-media-kanakmalpani
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: fal-ai-media
Source: https://github.com/KanakMalpani/General-Private-Skills/tree/main/skills/fal-ai-media
Command: npx skills add https://github.com/KanakMalpani/General-Private-Skills --skill fal-ai-media-kanakmalpani

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Unified media generation across images, videos, and audio using fal.ai MCP, enabling consistent, model-driven outputs without switching tools.

Core Features & Use Cases

  • Image generation from text prompts using Nano Banana models.
  • Video generation from prompts or inputs using Seedance, Kling, and Veo 3.
  • Audio generation including text-to-speech with CSM-1B and ThinkSound workflows, plus sample code for ElevenLabs integration.
  • Video-to-audio extraction and synchronization for multimedia projects.

Quick Start

Generate a 5-second video of a futuristic city at dusk using fal.ai MCP.

Frequently Asked Questions about fal-ai-media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate video and audio from text prompts using fal.ai?▼

Generate video and audio from text using fal.ai MCP by configuring the server and utilizing provided tools to access models like Seedance, Kling, Veo 3, and CSM-1B for multimedia synthesis.

What is the best way to create unified media assets without switching tools?▼

The best way to create unified media assets is using fal.ai MCP for model-driven generation across images, video, and audio, enabling consistent outputs without switching tools during rapid multimedia production.

Does fal.ai MCP support text-to-speech synthesis and video-to-audio extraction?▼

Yes, fal.ai MCP supports text-to-speech synthesis using CSM-1B and ThinkSound workflows, alongside video-to-audio extraction and synchronization for comprehensive multimedia projects.

Can I use ElevenLabs integration for audio generation with fal.ai MCP?▼

Yes, you can use ElevenLabs integration for audio generation; the Skill includes sample code for ElevenLabs workflows alongside native CSM-1B text-to-speech synthesis capabilities.

How do I start generating an image from text using Nano Banana models?▼

Start generating images from text by configuring your fal.ai MCP server, then use the supplied MCP tools like search and generate to execute text-to-image generation with Nano Banana models.