Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

williamlimasilva/phoenix-evals

Name: phoenix-evals
Author: williamlimasilva

skills/phoenix-evals/SKILL.md

npx skillsauth add williamlimasilva/.copilot phoenix-evals

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

| Task | Files | | ---- | ----- | | Setup | setup-python, setup-typescript | | Decide what to evaluate | evaluators-overview | | Choose a judge model | fundamentals-model-selection | | Use pre-built evaluators | evaluators-pre-built | | Build code evaluator | evaluators-code-python, evaluators-code-typescript | | Build LLM evaluator | evaluators-llm-python, evaluators-llm-typescript, evaluators-custom-templates | | Batch evaluate DataFrame | evaluate-dataframe-python | | Run experiment | experiments-running-python, experiments-running-typescript | | Create dataset | experiments-datasets-python, experiments-datasets-typescript | | Generate synthetic data | experiments-synthetic-python, experiments-synthetic-typescript | | Validate evaluator accuracy | validation, validation-evaluators-python, validation-evaluators-typescript | | Sample traces for review | observe-sampling-python, observe-sampling-typescript | | Analyze errors | error-analysis, error-analysis-multi-turn, axial-coding | | RAG evals | evaluators-rag | | Avoid common mistakes | common-mistakes-python, fundamentals-anti-patterns | | Production | production-overview, production-guardrails, production-continuous |

Workflows

Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview

Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Production: production-overview → production-guardrails → production-continuous

Reference Categories

| Prefix | Description | | ------ | ----------- | | fundamentals-* | Types, scores, anti-patterns | | observe-* | Tracing, sampling | | error-analysis-* | Finding failures | | axial-coding-* | Categorizing failures | | evaluators-* | Code, LLM, RAG evaluators | | experiments-* | Datasets, running experiments | | validation-* | Validating evaluator accuracy against human labels | | production-* | CI/CD, monitoring |

Key Principles

| Principle | Action | | --------- | ------ | | Error analysis first | Can't automate what you haven't observed | | Custom > generic | Build from your failures | | Code first | Deterministic before LLM | | Validate judges | >80% TPR/TNR | | Binary > Likert | Pass/fail, not 1-5 |

williamlimasilva/phoenix-evals

skills/phoenix-evals/SKILL.md

Build and run evaluators for AI/LLM applications using Phoenix.

development

Updated May 5, 2026

$ install --global

skillsauth

npx skillsauth add williamlimasilva/.copilot phoenix-evals

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security scan pending...

This skill is queued for security scanning. Results will appear when the scan completes.

SKILL.md

name:: phoenix-evals
description:: Build and run evaluators for AI/LLM applications using Phoenix.
license:: Apache-2.0
compatibility:: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
author:: [email protected]
version:: 1.0.0
languages:: Python, TypeScript

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

Workflows

Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview

Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Production: production-overview → production-guardrails → production-continuous

Reference Categories

Key Principles

Related Skills

williamlimasilva/workshop-create

tools

VerifiedTrustedCommunity

Create a new workshop or use an existing directory as one. Handles two paths: (A) use an existing local directory the operator points at, or (B) create a new private GitHub repo in the signed-in account. Never creates a repo inside another repo.

SKILL.mdUpdated Jul 22, 2026

williamlimasilva/workshop-create

williamlimasilva/vcpkg

development

VerifiedTrustedCommunity

Guide for setting up vcpkg in C++ projects, managing dependency versions, and cross-compiling. Covers manifest initialization, CMake and Visual Studio integration, classic-to-manifest migration, version pinning, baselines, overrides, triplets, and cross-compilation. Use when a user is working with vcpkg project setup, installation, version management, or cross-platform builds. For specialized tasks, additional references cover custom registries and overlay ports (references/registries.md), CI/CD and binary caching (references/ci.md), and troubleshooting and dependency lifecycle (references/troubleshooting.md).

SKILL.mdUpdated Jul 22, 2026

williamlimasilva/vcpkg

williamlimasilva/signal-write

testing

VerifiedTrustedCommunity

Emit structured agent signals — hands-up, blocked, done, checkpoint, partnership. Signals are written as JSON to .signals/ for dashboard consumption and noted in the journal for persistence.

SKILL.mdUpdated Jul 22, 2026

williamlimasilva/signal-write

williamlimasilva/markstream-install

development

VerifiedTrustedCommunity

Install and configure Markstream streaming Markdown renderers for Vue, React, Svelte, Angular, Nuxt, and Vue 2 applications. Use for package selection, minimal peer dependencies, CSS order, SSR boundaries, streaming mode, and renderer setup.

SKILL.mdUpdated Jul 22, 2026

williamlimasilva/markstream-install

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/williamlimasilva/.copilot.git

# Copy into Claude Code skills folder (global)
cp -r .copilot/skills/phoenix-evals ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

williamlimasilva/.copilot

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT