nemo-curator

Curate multimodal LLM training data with GPU-accelerated deduplication and filtering.

Updated Apr 23, 2026
One-click install
npx skills add https://github.com/Rawgrowth-Consulting/rawclaw-agent --skill nemo-curator-rawgrowth-consulting
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/Rawgrowth-Consulting/rawclaw-agent/tree/main/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/Rawgrowth-Consulting/rawclaw-agent --skill nemo-curator-rawgrowth-consulting

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Curates high-quality, multimodal training data for LLMs by accelerating GPU-based data processing, deduplication, and quality filtering.

Core Features & Use Cases

  • GPU-accelerated curation across text, image, video, and audio
  • Fuzzy, exact, and semantic deduplication
  • PII redaction and NSFW detection
  • Scales with RAPIDS for large-scale dataset preparation

Use cases include preparing large language model training data, cleaning web data, and deduplicating large corpora.

Quick Start

Launch the Nemo Curator pipeline to begin GPU-accelerated data curation on your training dataset.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I deduplicate large multimodal datasets for LLM training?▼

To deduplicate large multimodal datasets for LLM training, you can apply fuzzy, exact, and semantic deduplication. This process scales across text, image, video, and audio formats using GPU-accelerated RAPIDS environments.

Can I redact PII and detect NSFW content during data curation?▼

Yes, you can redact PII and detect NSFW content during data curation. These quality filtering steps are integrated directly into the GPU-accelerated pipeline to clean web data and prepare training corpora.

Does GPU-accelerated data curation work for text, image, video, and audio?▼

GPU-accelerated data curation works across text, image, video, and audio formats. It leverages RAPIDS multi-GPU scaling to process large-scale, multimodal datasets efficiently for LLM preparation.

What is the best way to clean web data for large language model training?▼

The best way to clean web data for large language model training is using a GPU-accelerated curation workflow. This approach combines deduplication, PII redaction, NSFW detection, and quality filtering to ensure high-quality inputs.

Do I need RAPIDS to scale data curation across multiple GPUs?▼

Yes, you need RAPIDS to scale data curation across multiple GPUs. RAPIDS provides the necessary multi-GPU scaling framework to handle large-scale dataset preparation and accelerate the curation pipeline.