llm-architect

Design production-grade LLM systems for scalable, safe deployment.

Updated Apr 27, 2026
One-click install
npx skills add https://github.com/Tnemo65/template --skill llm-architect-tnemo65
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: llm-architect
Source: https://github.com/Tnemo65/template/tree/main/.cursor/skills/11-system-design/llm-architect
Command: npx skills add https://github.com/Tnemo65/template --skill llm-architect-tnemo65

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Guides teams in designing, implementing, and operating production-grade LLM systems with strong performance, safety, and scalability.

Core Features & Use Cases

  • Architecture planning for serving patterns, multi-model routing, and monitoring
  • Fine-tuning, RAG integration, and cost-aware serving optimizations
  • Safety, compliance, and auditing through checklists and governance
  • End-to-end workflow from design to validation and deployment

Quick Start

Provide a scalable LLM system blueprint based on your current stack and latency targets.

Frequently Asked Questions about llm-architect

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable LLM serving architecture for enterprise workloads?▼

Design scalable LLM serving architecture by defining serving patterns, multi-model routing, and monitoring workflows to achieve target latency and throughput across enterprise environments.

What's the best way to integrate RAG into a production LLM system?▼

Integrate RAG into a production LLM system through structured integration workflows that combine retrieval pipelines with cost-aware serving optimizations to maintain performance and accuracy.

How do I set up monitoring and safety protocols for LLM deployment?▼

Set up LLM deployment monitoring and safety protocols by applying architecture checklists, governance auditing, and compliance validation to ensure safe and observable production operations.

Can I use multi-model orchestration to route requests across different LLMs?▼

Multi-model orchestration routes requests across different LLMs by applying architecture planning patterns that balance cost, latency, and performance targets for enterprise serving environments.

What do I need to plan before fine-tuning and deploying an LLM system?▼

Plan LLM system fine-tuning and deployment by preparing architecture blueprints, defining performance targets, and establishing safety checklists to validate end-to-end production workflows.