nemo-curator

Accelerate multi-modal LLM training data curation with GPU-accelerated pipelines.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/wwwillott/jobnimbus --skill nemo-curator-wwwillott
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: nemo-curator
Source: https://github.com/wwwillott/jobnimbus/tree/main/optional-skills/mlops/nemo-curator
Command: npx skills add https://github.com/wwwillott/jobnimbus --skill nemo-curator-wwwillott

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

NeMo Curator accelerates and scales the creation of high-quality, multi-modal training data for LLMs, reducing manual curation effort and improving dataset fidelity.

Core Features & Use Cases

  • GPU-accelerated quality filtering, deduplication, PII redaction, NSFW detection, and multilingual content assessment across text, image, audio, and video modalities
  • Multi-stage pipelines combining exact, fuzzy, and semantic deduplication with classifier-based filtering for production-grade data
  • Scalable, distributed workflows designed to prep large datasets like RedPajama v2 or The Pile

Quick Start

Instruct NeMo Curator to assemble a high-quality, multi-modal training dataset by applying deduplication, quality filtering, PII redaction, NSFW checks, and multi-GPU scaling.

Frequently Asked Questions about nemo-curator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I prepare large-scale multimodal training data for LLMs?▼

You prepare large-scale multimodal training data by applying GPU-accelerated pipelines for deduplication, quality filtering, and PII redaction across text, image, audio, and video modalities to yield high-fidelity datasets.

What is GPU-accelerated data curation for LLM training?▼

GPU-accelerated data curation uses distributed processing and frameworks like NVIDIA RAPIDS to optimize and scale the creation of high-quality training datasets, significantly reducing manual effort and processing time.

Does NeMo Curator support multilingual content assessment and NSFW detection?▼

Yes, NeMo Curator supports multilingual content assessment and NSFW detection as part of its modular, classifier-based filtering stages designed to ensure dataset quality and safety across modalities.

How do I perform exact and fuzzy deduplication on massive text datasets?▼

You perform exact and fuzzy deduplication on massive datasets by using multi-stage GPU-accelerated pipelines that combine semantic deduplication with classifier-based filtering for production-grade data curation.

Do I need NVIDIA GPUs to run distributed data curation workflows?▼

Yes, you need NVIDIA GPUs because the distributed data curation workflows rely on GPU-acceleration via NVIDIA RAPIDS and NeMo Curator frameworks to scale processing for large datasets like RedPajama v2 or The Pile.