Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

mediar-ai/packages/skills/skills/evaluation

Name: packages/skills/skills/evaluation
Author: mediar-ai

packages/skills/skills/evaluation/SKILL.md

npx skillsauth add mediar-ai/skillhubz packages/skills/skills/evaluation

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Evaluation Methods for Agent Systems

Evaluate agent systems with multi-dimensional rubrics accounting for non-determinism and multiple valid paths.

Prerequisites

Understanding of agent architectures
Familiarity with evaluation metrics

Instructions

Key Insight: Performance Drivers

Research found three factors explain 95% of performance variance:

| Factor | Variance | Implication | |--------|----------|-------------| | Token usage | 80% | More tokens = better performance | | Tool calls | ~10% | More exploration helps | | Model choice | ~5% | Better models multiply efficiency |

Multi-Dimensional Rubric

| Dimension | Measures | |-----------|----------| | Factual accuracy | Claims match ground truth | | Completeness | Output covers requested aspects | | Citation accuracy | Citations match sources | | Source quality | Uses appropriate primary sources | | Tool efficiency | Right tools, reasonable count |

Test Set Design

Complexity Stratification:

Simple: Single tool call
Medium: Multiple tool calls
Complex: Many calls, significant ambiguity
Very complex: Extended interaction, deep reasoning

Evaluation Methodologies

LLM-as-Judge: Scalable, consistent judgments.

Provide task, output, ground truth, evaluation scale
Request structured judgment

Human Evaluation: Catches what automation misses.

Edge cases, system failures, subtle biases

End-State Evaluation: For agents that mutate state.

Focus on final state matching expectations

Guidelines

Use multi-dimensional rubrics, not single metrics
Evaluate outcomes, not specific execution paths
Cover complexity levels from simple to complex
Test with realistic context sizes
Supplement LLM evaluation with human review
Track metrics over time for trend detection

Notes

Agents may take different valid paths to goals
Model upgrades beat token increases for performance
High absolute agreement matters less than systematic patterns

Source: muratcankoylan/Agent-Skills-for-Context-Engineering

mediar-ai/packages/skills/skills/evaluation

packages/skills/skills/evaluation/SKILL.md

# Evaluation Methods for Agent Systems Evaluate agent systems with multi-dimensional rubrics accounting for non-determinism and multiple valid paths. ## Prerequisites - Understanding of agent architectures - Familiarity with evaluation metrics ## Instructions ### Key Insight: Performance Drivers Research found three factors explain 95% of performance variance: | Factor | Variance | Implication | |--------|----------|-------------| | Token usage | 80% | More tokens = better performance | |

4 stars

data-ai

Updated Apr 18, 2026

$ install --global

skillsauth

npx skillsauth add mediar-ai/skillhubz packages/skills/skills/evaluation

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 18, 2026, 6:50 AM372.7s2 files scanned

SKILL.md

Evaluation Methods for Agent Systems

Evaluate agent systems with multi-dimensional rubrics accounting for non-determinism and multiple valid paths.

Prerequisites

Understanding of agent architectures
Familiarity with evaluation metrics

Instructions

Key Insight: Performance Drivers

Research found three factors explain 95% of performance variance:

Multi-Dimensional Rubric

Test Set Design

Complexity Stratification:

Simple: Single tool call
Medium: Multiple tool calls
Complex: Many calls, significant ambiguity
Very complex: Extended interaction, deep reasoning

Evaluation Methodologies

LLM-as-Judge: Scalable, consistent judgments.

Provide task, output, ground truth, evaluation scale
Request structured judgment

Human Evaluation: Catches what automation misses.

Edge cases, system failures, subtle biases

End-State Evaluation: For agents that mutate state.

Focus on final state matching expectations

Guidelines

Use multi-dimensional rubrics, not single metrics
Evaluate outcomes, not specific execution paths
Cover complexity levels from simple to complex
Test with realistic context sizes
Supplement LLM evaluation with human review
Track metrics over time for trend detection

Notes

Agents may take different valid paths to goals
Model upgrades beat token increases for performance
High absolute agreement matters less than systematic patterns

Source: muratcankoylan/Agent-Skills-for-Context-Engineering

Related Skills

mediar-ai/tui-ui

tools

VerifiedTrustedCommunity

Design web-like user interfaces in the terminal and inside tmux with a cell-grid Canvas, CSS-like box model, flexbox/grid layout, and 15 reusable widgets such as Panel, Table, Card, ProgressBar, Meter, Tabs, Tree, Badge, Banner, and a braille line chart. Use when an agent needs a dashboard, panel, table, status page, TUI layout, tmux dashboard, screenshot-driven CLI/TUI replica, ANSI frame, truecolor render, pyte PNG screenshot smoke test, wide-character alignment, or a new terminal widget.

7SKILL.mdUpdated Jul 11, 2026

mediar-ai/drive-tui

tools

VerifiedTrustedCommunity

Drive interactive terminal (TUI) programs — CLIs, REPLs, installers, menu apps, agent CLIs, and editors like vim — through a PTY, reading semantic screen snapshots. A pattern library classifies a screen (REPL, menu, pager, fzf search, confirm dialog, form, spinner, wizard) and drives it with a ready recipe. Use when a program expects a live terminal (arrow-key menus, prompts, spinners, password fields, curses UIs), or when a piped command hangs or prints nothing.

7SKILL.mdUpdated Jul 11, 2026

mediar-ai/cmd-art

tools

VerifiedTrustedCommunity

Design and render terminal/CMD visual effects and ASCII art from a one-line request via the pluggable `fx` engine (18 hot-swappable, themeable effects plus scripted shows). Effects include donut, matrix rain, plasma, fire, a spinning 3D ball, Game of Life, wireframe cube, 3D text banners, rainbow/lolcat gradient text, starfield, tunnel, fireworks, image-to-ASCII, and more. Use when the request is for a terminal animation, ANSI/CLI art, or a new console effect. Pure Python stdlib; truecolor.

7SKILL.mdUpdated Jul 11, 2026

mediar-ai/packages/skills/skills/x-twitter-scraper

tools

VerifiedTrustedCommunity

# X Twitter Scraper Use Xquik for X/Twitter tweet search, user lookup, profile tweets, follower export, media download, monitors, webhooks, posting workflows, and MCP-backed API exploration. ## Prerequisites - A Xquik API key in `XQUIK_API_KEY`. - Internet access to `https://xquik.com/api/v1`, `https://xquik.com/mcp`, and `https://docs.xquik.com`. - A clear user request that identifies the target tweets, users, accounts, keywords, media, monitor, webhook, or write action. ## Source Truth -

6SKILL.mdUpdated May 31, 2026

mediar-ai/packages/skills/skills/x-twitter-scraper

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/mediar-ai/skillhubz.git

# Copy into Claude Code skills folder (global)
cp -r skillhubz/packages/skills/skills/evaluation ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

mediar-ai/skillhubz

4 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT