ZJUNLP
Official@zjunlp · China
Knowledge Engine Lab: A NLP & KG Group of Zhejiang University
Agent Skills by ZJUNLP
Showing 222 vetted skills indexed across 3 GitHub repositories.
auto-verify
Stress-test research claims by swapping method, dataset, and model variants.
impact-check
Assess the importance and reach of a research idea before committing effort.
notify
Drafts and dispatches research-progress briefings through user-configured notification services.
mechanism-behavior-discovery
Surfaces novel falsifiable behavioral phenomena in LLMs for downstream mechanistic investigation.
training-check
Monitors WandB training metrics periodically to detect NaN, divergence, and stalled runs.
mechanism-skills
Route mechanistic interpretability questions to eleven method families for localizing internal model components.
result-to-claim
Evaluates experiment results against intended claims using an external LLM reviewer and routes next actions.
hypothesis-batch
Generates and refines batches of research hypotheses through a multi-phase automated pipeline.
auto
Orchestrates autonomous research pipelines from claim generation through experiments to verification and iteration.
research-refine-pipeline
Chains method refinement and experiment planning into one end-to-end research proposal workflow.
experiment-audit
Audits per-claim experimental methodology integrity using cross-model LLM review.
experiment-plan
Converts a refined research proposal into a claim-driven experiment roadmap with run order and budgets.
idea-creator
Generate, validate, and rank research ideas with pilot experiments for a given direction.
experiment-queue
Orchestrates batched ML experiments on SSH GPU servers with OOM retry and wave scheduling.
mhistory
Generates a chronological research development-history article from database retrieval and web search.
mechanic-db-search
Retrieves academic papers from interpretability and cross-disciplinary databases via a cloud search service.
monitor-experiment
Monitor remote GPU experiments, collect results, and finalize cost manifests over SSH.
research-refine
Refines vague research directions into concrete method proposals via iterative external LLM review.
data-rule
Defines dataset provenance, split, label, and sample-size constraints for experiments.
ablation-planner
Designs and runs ablation studies for ML experiments using an external LLM reviewer.
auto-claim
Orchestrates the claim-stage pipeline producing research proposals and experiment plans.
mechanism-audit
Audit mechanistic interpretability experiment rigor per claim using cross-model LLM review.
analyze-results
Analyze ML experiment results and generate comparison tables with statistical insights.
paper-figure
Generate publication-quality matplotlib figures and LaTeX tables from experiment data files.
Frequently Asked Questions About ZJUNLP
FAQPage SchemaWhat specific tasks can I perform using ZJUNLP's simulated environment skills?▼
You can execute complex object manipulation tasks, including locating, heating, cooling, cleaning, and storing items within ALFWorld and ScienceWorld environments, as well as conducting scientific experiments like circuit building and substance mixing.
Which personas benefit most from these technical capabilities?▼
Researchers and developers focused on embodied cognition, decision-making benchmarks, and multi-turn data analysis will find these capabilities most relevant for testing and refining complex reasoning systems.
What are the primary data dependencies for the financial and clinical analysis skills?▼
These skills require structured datasets, specifically SQLite databases containing MIMIC-IV patient records or SEC 10-K financial filings, to perform metric extraction and report generation.