packages/skills/skills/evaluation/SKILL.md
# Evaluation Methods for Agent Systems Evaluate agent systems with multi-dimensional rubrics accounting for non-determinism and multiple valid paths. ## Prerequisites - Understanding of agent architectures - Familiarity with evaluation metrics ## Instructions ### Key Insight: Performance Drivers Research found three factors explain 95% of performance variance: | Factor | Variance | Implication | |--------|----------|-------------| | Token usage | 80% | More tokens = better performance | |
npx skillsauth add mediar-ai/skillhubz packages/skills/skills/evaluationInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Evaluate agent systems with multi-dimensional rubrics accounting for non-determinism and multiple valid paths.
Research found three factors explain 95% of performance variance:
| Factor | Variance | Implication | |--------|----------|-------------| | Token usage | 80% | More tokens = better performance | | Tool calls | ~10% | More exploration helps | | Model choice | ~5% | Better models multiply efficiency |
| Dimension | Measures | |-----------|----------| | Factual accuracy | Claims match ground truth | | Completeness | Output covers requested aspects | | Citation accuracy | Citations match sources | | Source quality | Uses appropriate primary sources | | Tool efficiency | Right tools, reasonable count |
Complexity Stratification:
LLM-as-Judge: Scalable, consistent judgments.
Human Evaluation: Catches what automation misses.
End-State Evaluation: For agents that mutate state.
Source: muratcankoylan/Agent-Skills-for-Context-Engineering
tools
Design web-like user interfaces in the terminal and inside tmux with a cell-grid Canvas, CSS-like box model, flexbox/grid layout, and 15 reusable widgets such as Panel, Table, Card, ProgressBar, Meter, Tabs, Tree, Badge, Banner, and a braille line chart. Use when an agent needs a dashboard, panel, table, status page, TUI layout, tmux dashboard, screenshot-driven CLI/TUI replica, ANSI frame, truecolor render, pyte PNG screenshot smoke test, wide-character alignment, or a new terminal widget.
tools
Drive interactive terminal (TUI) programs — CLIs, REPLs, installers, menu apps, agent CLIs, and editors like vim — through a PTY, reading semantic screen snapshots. A pattern library classifies a screen (REPL, menu, pager, fzf search, confirm dialog, form, spinner, wizard) and drives it with a ready recipe. Use when a program expects a live terminal (arrow-key menus, prompts, spinners, password fields, curses UIs), or when a piped command hangs or prints nothing.
tools
Design and render terminal/CMD visual effects and ASCII art from a one-line request via the pluggable `fx` engine (18 hot-swappable, themeable effects plus scripted shows). Effects include donut, matrix rain, plasma, fire, a spinning 3D ball, Game of Life, wireframe cube, 3D text banners, rainbow/lolcat gradient text, starfield, tunnel, fireworks, image-to-ASCII, and more. Use when the request is for a terminal animation, ANSI/CLI art, or a new console effect. Pure Python stdlib; truecolor.
tools
# X Twitter Scraper Use Xquik for X/Twitter tweet search, user lookup, profile tweets, follower export, media download, monitors, webhooks, posting workflows, and MCP-backed API exploration. ## Prerequisites - A Xquik API key in `XQUIK_API_KEY`. - Internet access to `https://xquik.com/api/v1`, `https://xquik.com/mcp`, and `https://docs.xquik.com`. - A clear user request that identifies the target tweets, users, accounts, keywords, media, monitor, webhook, or write action. ## Source Truth -