What problem does it solve? It lets an AI agent convert text into natural-sounding speech and transcribe recorded audio into text through Google's Gemini voice models, without the user needing to learn the underlying API details. ## Core Features & Use Cases - Text-to-Speech: Gemini 3.1 Flash TTS turns text into speech with 30 voices, natural-language delivery control, and two-speaker dialogue. - Speech-to-Text: Gemini 3.5 Transcribe converts recorded speech into text with speaker labels and word-level timestamps across 85+ locales. - Guided API Key Setup: Walks the agent through acquiring and saving a Pixazo API key once, then reuses it for all future calls. - Use Case: Ask your agent to read a blog post aloud as an MP3, or transcribe a recorded meeting with speaker labels and timestamps. ## Quick Start Ask your agent to convert this paragraph into spoken audio using Gemini Voice and give me the audio file URL.