llava

Answer questions about images using CLIP-based vision encoding and Vicuna/LLaMA language models.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/Hermesagents/hermes-agents --skill llava-hermesagents
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: llava
Source: https://github.com/Hermesagents/hermes-agents/tree/main/skills/mlops/models/llava
Command: npx skills add https://github.com/Hermesagents/hermes-agents --skill llava-hermesagents

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides a unified vision-language capability that lets users ask questions about images and receive coherent, context-aware responses.

Core Features & Use Cases

  • Multimodal image conversation: discuss visuals across multiple turns.
  • Visual question answering (VQA): answer questions about objects, scenes, and actions in images.
  • Image captioning and description: generate natural language descriptions of visual content.
  • Instruction following with visual context: perform tasks based on images and prompts.

Quick Start

Upload an image and start a multimodal chat to get descriptive, analytical, or instruction-driven responses.

Frequently Asked Questions about llava

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How does visual question answering work for image chat?▼

Visual question answering for image chat works by applying CLIP-based vision encoding with Vicuna or LLaMA-style language models to answer questions about objects, scenes, and actions in images. It supports multi-turn dialogs with visual context.

What is multimodal instruction-following for visual content?▼

Multimodal instruction-following for visual content is a vision-language task where you upload an image and provide prompts to receive coherent, context-aware responses. It lets you perform tasks based on images and instructions across iterative conversations.

Can I use Vicuna and LLaMA models for multi-turn image conversations?▼

Yes, you can use Vicuna and LLaMA-style language models for multi-turn image conversations. The system relies on CLIP-based vision encoding combined with these language models to support iterative conversations with visual context.

How do I start an image chat to generate descriptions of visual content?▼

To start an image chat and generate descriptions of visual content, simply upload an image and begin a multimodal chat. You will receive descriptive, analytical, or instruction-driven responses about the visual content.

What is the best way to perform visual instruction-following across multiple turns?▼

The best way to perform visual instruction-following across multiple turns is to use a unified vision-language capability that maintains visual context. This approach lets you ask questions and give prompts about images over iterative conversations.

Does this image chat approach support CLIP-based vision encoding with language models?▼

Yes, this image chat approach supports CLIP-based vision encoding integrated with Vicuna or LLaMA-style language models. This combination enables context-aware responses for visual question answering and image captioning tasks.