jarvis-voice

Generate local metallic text-to-speech with sherpa-onnx and customizable effects.

Updated Feb 17, 2026
One-click install
npx skills add https://github.com/Qcasares/saas-app --skill jarvis-voice-qcasares
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: jarvis-voice
Source: https://github.com/Qcasares/saas-app/tree/main/skills/jarvis-voice
Command: npx skills add https://github.com/Qcasares/saas-app --skill jarvis-voice-qcasares

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires sherpa-onnx, ffmpeg, aplay, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a local, metallic text-to-speech (TTS) voice for AI agents, enhancing user engagement and accessibility through visual transcripts and voice effects.

Core Features & Use Cases

  • Metallic Voice: Local speech synthesis with sherpa-onnx, avoiding cloud latency.
  • Customizable Effects: Apply effects like flanger, echo, and pitch shift for unique voices.
  • Visual Transcripts: Differentiate spoken text visually in webchat with purple italic styling.
  • Fast Playback: Enable 2x speed for quick communication.
  • Use Case: Improve the user experience for AI-powered chatbots and assistants by adding a robotic, distinctive voice that can be customized with various effects.

Quick Start

Use the jarvis-voice skill to generate a speech output for the text 'Hello, I am your AI assistant.'.

Frequently Asked Questions about jarvis-voice

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I add local text-to-speech to an AI agent without cloud APIs?▼

Local text-to-speech for AI agents can be generated using sherpa-onnx to synthesize speech without relying on cloud APIs. This approach processes audio locally to avoid cloud latency while providing customizable voice effects and visual transcripts.

What voice effects can I apply to local TTS output for chatbots?▼

Voice effects for local TTS output include flanger, echo, and pitch shift to create unique metallic voices. These effects allow you to customize speech synthesis, enhancing user engagement for AI-powered chatbots and assistants.

Does local TTS processing require sherpa-onnx and ffmpeg?▼

Local TTS processing requires sherpa-onnx, ffmpeg, and aplay to function properly without cloud dependencies. These tools handle speech synthesis, audio processing, and playback respectively, ensuring the local metallic voice generation works as intended.

Can I display visual transcripts alongside speech output in webchat?▼

Visual transcripts can be displayed alongside speech output in webchat using purple italic styling to differentiate spoken text. This feature improves accessibility and user engagement by providing a clear visual representation of the synthesized speech.

How do I enable fast playback for AI assistant voice output?▼

Fast playback for AI assistant voice output can be enabled by applying a 2x speed setting to the synthesized speech. This allows for quick communication, reducing the time users spend listening to responses from the local metallic TTS system.