What problem does it solve? Reviewing sdd-lite evaluation trends, regressions, and token usage requires aggregating scattered run evidence without launching another expensive model-backed test. This Skill rebuilds campaign reports and local history from existing evidence, or compares a candidate run against an explicitly named baseline. ## Core Features & Use Cases - Campaign Report Rebuild: Regenerate evaluation reports for a workspace and campaign using the sddl_eval.py report command. - Explicit Baseline Comparison: Compare a candidate run against a named baseline run, flagging comparisons as confounded when provider, case, project hash, or model differ. - Verdict and Cost Summaries: Summarize verdict deltas, new/persisting/resolved findings, and token or cost changes while keeping raw transcripts out of the human summary. - Use Case: After a week of sdd-lite evaluation runs, rebuild the campaign report to spot regressions and compare yesterday's candidate run against last week's baseline to explain token usage differences. ## Quick Start Ask the assistant to rebuild the sdd-lite evaluation report for your workspace and campaign, or to compare a specific candidate run against a named baseline run.