Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

dwmkerr/anthropic-evaluations

Name: anthropic-evaluations
Author: dwmkerr

plugins/toolkit/skills/anthropic-evaluations/SKILL.md

npx skillsauth add dwmkerr/claude-toolkit anthropic-evaluations

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Anthropic Evaluations

Build rigorous evaluations for AI agents using Anthropic's proven patterns.

Quick Reference

You MUST read the reference files for detailed guidance:

Grader Types - Code-based, model-based, human graders
Agent Type Patterns - Coding, conversational, research, computer use
Roadmap - Steps 0-8 for building evals from scratch
Frameworks - Harbor, Promptfoo, Braintrust, etc.

YAML Templates:

coding-agent-eval.yaml - Coding agent template
conversational-agent-eval.yaml - Support agent template

Annotated Examples:

Example: Coding Agent - Auth bypass fix walkthrough
Example: Conversational - Refund handling walkthrough

Core Definitions

| Term | Definition | |------|------------| | Task | Single test with defined inputs and success criteria | | Trial | One attempt at a task (run multiple for consistency) | | Grader | Logic that scores agent performance; tasks can have multiple | | Transcript | Complete record of a trial (outputs, tool calls, reasoning) | | Outcome | Final state in environment (not just what agent said) | | Evaluation harness | Infrastructure that runs evals end-to-end | | Agent harness | System enabling model to act as agent (scaffold) | | Evaluation suite | Collection of tasks measuring specific capabilities |

Grader Types (Quick Reference)

| Type | Methods | Best For | |------|---------|----------| | Code-based | String match, unit tests, static analysis, state checks | Fast, cheap, objective verification | | Model-based | Rubric scoring, assertions, pairwise comparison | Nuanced, open-ended tasks | | Human | SME review, A/B testing, spot-check sampling | Gold standard calibration |

See Grader Types for detailed comparison.

Capability vs Regression Evals

| Type | Question | Target Pass Rate | |------|----------|------------------| | Capability | "What can this agent do well?" | Start low, hill-climb | | Regression | "Does it still handle what it used to?" | Near 100% |

Capability evals with high pass rates "graduate" to regression suites.

Non-Determinism Metrics

| Metric | Measures | Use When | |--------|----------|----------| | pass@k | At least 1 success in k attempts | One success matters (coding) | | pass^k | All k attempts succeed | Consistency essential (customer-facing) |

Example: 75% per-trial success rate

pass@3 ≈ 98% (likely to get at least one)
pass^3 ≈ 42% (0.75³ all succeed)

Tracked Metrics

tracked_metrics:
  - type: transcript
    metrics: [n_turns, n_toolcalls, n_total_tokens]
  - type: latency
    metrics: [time_to_first_token, output_tokens_per_sec, time_to_last_token]

Attribution

Based on Demystifying evals for AI agents by Anthropic (January 2026).

dwmkerr/anthropic-evaluations

plugins/toolkit/skills/anthropic-evaluations/SKILL.md

This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also suggest when building coding agents, conversational agents, or research agents that need quality assurance.

12 stars

development

Updated Apr 25, 2026

$ install --global

skillsauth

npx skillsauth add dwmkerr/claude-toolkit anthropic-evaluations

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 25, 2026, 3:39 PM121.5s9 files scanned

SKILL.md

name:: anthropic-evaluations
description:: This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also suggest when building coding agents, conversational agents, or research agents that need quality assurance.
allowed-tools:: Read, Grep

Anthropic Evaluations

Build rigorous evaluations for AI agents using Anthropic's proven patterns.

Quick Reference

You MUST read the reference files for detailed guidance:

Grader Types - Code-based, model-based, human graders
Agent Type Patterns - Coding, conversational, research, computer use
Roadmap - Steps 0-8 for building evals from scratch
Frameworks - Harbor, Promptfoo, Braintrust, etc.

YAML Templates:

coding-agent-eval.yaml - Coding agent template
conversational-agent-eval.yaml - Support agent template

Annotated Examples:

Example: Coding Agent - Auth bypass fix walkthrough
Example: Conversational - Refund handling walkthrough

Core Definitions

Grader Types (Quick Reference)

See Grader Types for detailed comparison.

Capability vs Regression Evals

Capability evals with high pass rates "graduate" to regression suites.

Non-Determinism Metrics

Example: 75% per-trial success rate

pass@3 ≈ 98% (likely to get at least one)
pass^3 ≈ 42% (0.75³ all succeed)

Tracked Metrics

tracked_metrics:
  - type: transcript
    metrics: [n_turns, n_toolcalls, n_total_tokens]
  - type: latency
    metrics: [time_to_first_token, output_tokens_per_sec, time_to_last_token]

Attribution

Based on Demystifying evals for AI agents by Anthropic (January 2026).

Related Skills

dwmkerr/claude-code-skill-development

tools

VerifiedTrustedCommunity

This skill should be used when the user asks to "create a skill", "write a skill", "build a skill", or wants to add new capabilities to Claude Code. Use when developing SKILL.md files, organizing skill content, or improving existing skills. Do NOT use for plugin development, hook creation, agent creation, or slash command creation — those have dedicated skills.

12SKILL.mdUpdated Apr 25, 2026

dwmkerr/claude-code-skill-development

dwmkerr/shell-script-development

development

VerifiedTrustedCommunity

This skill should be used when the user asks to "create a bash script", "write a shell script", or mentions shell scripting conventions.

12SKILL.mdUpdated Apr 25, 2026

dwmkerr/shell-script-development

dwmkerr/research

development

VerifiedTrustedCommunity

Deep research into technical solutions by searching the web, examining GitHub repos, and gathering evidence. Use when the user explicitly says "use the research skill", "use a research agent", or asks for deep/thorough research into implementation options or technologies.

12SKILL.mdUpdated Apr 25, 2026

dwmkerr/release-please-development

tools

VerifiedTrustedCommunity

This skill should be used when the user asks to "set up release please", "configure automated releases", "manage version numbers", "add changelog automation", or mentions release-please, semantic versioning, or monorepo versioning.

12SKILL.mdUpdated Apr 25, 2026

dwmkerr/release-please-development

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/dwmkerr/claude-toolkit.git

# Copy into Claude Code skills folder (global)
cp -r claude-toolkit/plugins/toolkit/skills/anthropic-evaluations ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

dwmkerr/claude-toolkit

12 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT