aiml-nsfw-benchmark

Benchmark NSFW text classification across five categories using a standardized ISC template.

934|152|Updated Mar 1, 2026
One-click install
npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-nsfw-benchmark
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: aiml-nsfw-benchmark
Source: https://github.com/wuyoscar/ISC-Bench/tree/main/templates/aiml_nsfw
Command: npx skills add https://github.com/wuyoscar/ISC-Bench --skill aiml-nsfw-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The NSFW benchmarking template provides a standardized, reproducible framework to evaluate DistilBERT NSFW text classification across multiple explicit content categories, enabling researchers to measure recall, coverage, and robustness in a controlled setting.

Core Features & Use Cases

  • Standardized NSFW categories (sexual_explicit, sexual_suggestive, harassment_sexual, grooming, sextortion) and language variants to test cross-language capabilities.
  • Ready-to-use prompts and sample datasets with deterministic evaluation flows, plus guidance for extending categories and anchors.
  • Use cases include model evaluation in AI safety research, benchmark-driven fine-tuning, and validation of prompt-variant experiments.

Quick Start

Run the included prompts and scripts to evaluate your NSFW classifier on the provided five-category benchmark.

Frequently Asked Questions about aiml-nsfw-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark NSFW text classification across multiple categories?▼

You can benchmark NSFW text classification using a standardized ISC template to evaluate recall and category coverage across sexual_explicit, sexual_suggestive, harassment_sexual, grooming, and sextortion categories. The benchmark provides reproducible evaluation workflows with JSON samples and Python tooling.

What categories does the NSFW benchmark evaluate for AI safety research?▼

The NSFW benchmark evaluates five categories: sexual_explicit, sexual_suggestive, harassment_sexual, grooming, and sextortion. These standardized categories enable researchers to measure recall, coverage, and robustness in controlled AI safety evaluation settings.

Can I test cross-language NSFW detection capabilities with this benchmark?▼

Yes, the NSFW benchmark supports multi-language prompt sets and language variants to test cross-language NSFW detection capabilities. This allows evaluation of model performance across different languages for comprehensive AI safety validation.

How do I evaluate DistilBERT NSFW classifier performance with reproducible workflows?▼

You can evaluate DistilBERT NSFW classifier performance using the benchmark's deterministic evaluation flows with ready-to-use prompts and sample datasets. The included Python tooling and JSON samples ensure reproducible evaluation workflows for measuring recall and category coverage.

Does the NSFW benchmark support weak-anchor variant testing for prompt experiments?▼

Yes, the NSFW benchmark supports weak-anchor variants for prompt-variant experiments. This feature enables researchers to validate prompt-variant experiments and extend categories and anchors as needed for comprehensive model evaluation.

What's the best way to validate recall and category coverage in NSFW detection models?▼

The best way to validate recall and category coverage is using a standardized NSFW benchmark with reproducible evaluation workflows. The framework provides deterministic evaluation flows across five explicit content categories with JSON samples and Python tooling for measuring robustness.