Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

frank-luongt/skills/codex/databricks-mlflow-evaluation

Name: skills/codex/databricks-mlflow-evaluation
Author: frank-luongt

skills/codex/databricks-mlflow-evaluation/SKILL.md

npx skillsauth add frank-luongt/faos-skills-marketplace skills/codex/databricks-mlflow-evaluation

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

name: databricks-mlflow-evaluation

MLflow 3 GenAI Evaluation

Before Writing Any Code

Read GOTCHAS.md - 15+ common mistakes that cause failures
Read CRITICAL-interfaces.md - Exact API signatures and data schemas

End-to-End Workflows

Follow these workflows based on your goal. Each step indicates which reference files to read.

Workflow 1: First-Time Evaluation Setup

For users new to MLflow GenAI evaluation or setting up evaluation for a new agent.

| Step | Action | Reference Files | | ---- | --------------------------- | ---------------------------------------------------------------- | | 1 | Understand what to evaluate | user-journeys.md (Journey 0: Strategy) | | 2 | Learn API patterns | GOTCHAS.md + CRITICAL-interfaces.md | | 3 | Build initial dataset | patterns-datasets.md (Patterns 1-4) | | 4 | Choose/create scorers | patterns-scorers.md + CRITICAL-interfaces.md (built-in list) | | 5 | Run evaluation | patterns-evaluation.md (Patterns 1-3) |

Workflow 2: Production Trace -> Evaluation Dataset

For building evaluation datasets from production traces.

| Step | Action | Reference Files | | ---- | ----------------------------- | ------------------------------------------------ | | 1 | Search and filter traces | patterns-trace-analysis.md (MCP tools section) | | 2 | Analyze trace quality | patterns-trace-analysis.md (Patterns 1-7) | | 3 | Tag traces for inclusion | patterns-datasets.md (Patterns 16-17) | | 4 | Build dataset from traces | patterns-datasets.md (Patterns 6-7) | | 5 | Add expectations/ground truth | patterns-datasets.md (Pattern 2) |

Workflow 3: Performance Optimization

For debugging slow or expensive agent execution.

| Step | Action | Reference Files | | ---- | ----------------------------- | ---------------------------------------------------- | | 1 | Profile latency by span | patterns-trace-analysis.md (Patterns 4-6) | | 2 | Analyze token usage | patterns-trace-analysis.md (Pattern 9) | | 3 | Detect context issues | patterns-context-optimization.md (Section 5) | | 4 | Apply optimizations | patterns-context-optimization.md (Sections 1-4, 6) | | 5 | Re-evaluate to measure impact | patterns-evaluation.md (Pattern 6-7) |

Workflow 4: Regression Detection

For comparing agent versions and finding regressions.

| Step | Action | Reference Files | | ---- | ----------------------- | ------------------------------------------------ | | 1 | Establish baseline | patterns-evaluation.md (Pattern 4: named runs) | | 2 | Run current version | patterns-evaluation.md (Pattern 1) | | 3 | Compare metrics | patterns-evaluation.md (Patterns 6-7) | | 4 | Analyze failing traces | patterns-trace-analysis.md (Pattern 7) | | 5 | Debug specific failures | patterns-trace-analysis.md (Patterns 8-9) |

Workflow 5: Custom Scorer Development

For creating project-specific evaluation metrics.

| Step | Action | Reference Files | | ---- | --------------------------- | ----------------------------------------- | | 1 | Understand scorer interface | CRITICAL-interfaces.md (Scorer section) | | 2 | Choose scorer pattern | patterns-scorers.md (Patterns 4-11) | | 3 | For multi-agent scorers | patterns-scorers.md (Patterns 13-16) | | 4 | Test with evaluation | patterns-evaluation.md (Pattern 1) |

Reference Files Quick Lookup

| Reference | Purpose | When to Read | | ---------------------------------- | ------------------------ | ----------------------------------------- | | GOTCHAS.md | Common mistakes | Always read first before writing code | | CRITICAL-interfaces.md | API signatures, schemas | When writing any evaluation code | | patterns-evaluation.md | Running evals, comparing | When executing evaluations | | patterns-scorers.md | Custom scorer creation | When built-in scorers aren't enough | | patterns-datasets.md | Dataset building | When preparing evaluation data | | patterns-trace-analysis.md | Trace debugging | When analyzing agent behavior | | patterns-context-optimization.md | Token/latency fixes | When agent is slow or expensive | | user-journeys.md | High-level workflows | When starting a new evaluation project |

Critical API Facts

Use: mlflow.genai.evaluate() (NOT mlflow.evaluate())
Data format: {"inputs": {"query": "..."}} (nested structure required)
predict_fn: Receives **unpacked kwargs (not a dict)

See GOTCHAS.md for complete list.

frank-luongt/skills/codex/databricks-mlflow-evaluation

skills/codex/databricks-mlflow-evaluation/SKILL.md

--- name: databricks-mlflow-evaluation --- # MLflow 3 GenAI Evaluation ## Before Writing Any Code 1. **Read GOTCHAS.md** - 15+ common mistakes that cause failures 2. **Read CRITICAL-interfaces.md** - Exact API signatures and data schemas ## End-to-End Workflows Follow these workflows based on your goal. Each step indicates which reference files to read. ### Workflow 1: First-Time Evaluation Setup For users new to MLflow GenAI evalu

26 stars

development

Updated Jul 9, 2026

$ install --global

skillsauth

npx skillsauth add frank-luongt/faos-skills-marketplace skills/codex/databricks-mlflow-evaluation

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 21, 2026, 6:10 AM50.4s4 files scanned

SKILL.md

name: databricks-mlflow-evaluation

MLflow 3 GenAI Evaluation

Before Writing Any Code

Read GOTCHAS.md - 15+ common mistakes that cause failures
Read CRITICAL-interfaces.md - Exact API signatures and data schemas

End-to-End Workflows

Follow these workflows based on your goal. Each step indicates which reference files to read.

Workflow 1: First-Time Evaluation Setup

For users new to MLflow GenAI evaluation or setting up evaluation for a new agent.

Workflow 2: Production Trace -> Evaluation Dataset

For building evaluation datasets from production traces.

Workflow 3: Performance Optimization

For debugging slow or expensive agent execution.

Workflow 4: Regression Detection

For comparing agent versions and finding regressions.

Workflow 5: Custom Scorer Development

For creating project-specific evaluation metrics.

Reference Files Quick Lookup

Critical API Facts

Use: mlflow.genai.evaluate() (NOT mlflow.evaluate())
Data format: {"inputs": {"query": "..."}} (nested structure required)
predict_fn: Receives **unpacked kwargs (not a dict)

See GOTCHAS.md for complete list.

Related Skills

frank-luongt/skills/codex/grpo-rl-training

development

VerifiedTrustedCommunity

--- name: grpo-rl-training description: GRPO reinforcement learning training with TRL. Use when applying Group Relative Policy Optimization for reasoning and task-specific model training. --- # GRPO/RL Training with TRL Expert-level guidance for implementing Group Relative Policy Optimization (GRPO) using the Transformer Reinforcement Learning (TRL) library. This skill provides battle-tested patterns, critical insights, and production-r

26SKILL.mdUpdated Jul 9, 2026

frank-luongt/skills/codex/grpo-rl-training

frank-luongt/skills/codex/graphql-architect

tools

VerifiedTrustedCommunity

--- name: graphql-architect description: Master modern GraphQL with federation, performance optimization, --- ## Use this skill when - Working on graphql architect tasks or workflows - Needing guidance, best practices, or checklists for graphql architect ## Do not use this skill when - The task is unrelated to graphql architect - You need a different domain or tool outside this scope ## Instructions - Clarify goals, constraints, and

26SKILL.mdUpdated Jul 9, 2026

frank-luongt/skills/codex/graphql-architect

frank-luongt/skills/codex/grafana-dashboards

development

VerifiedTrustedCommunity

--- name: grafana-dashboards description: Create and manage production Grafana dashboards for real-time visualization of system and application metrics. Use when building monitoring dashboards, visualizing metrics, or creating operational observability interfaces. --- # Grafana Dashboards Create and manage production-ready Grafana dashboards for comprehensive system observability. ## Do not use this skill when - The task is unrelated

26SKILL.mdUpdated Jul 9, 2026

frank-luongt/skills/codex/grafana-dashboards

frank-luongt/skills/codex/gptq

development

VerifiedTrustedCommunity

--- name: gptq description: GPTQ post-training quantization for generative models. Use when quantizing large models to 4-bit with calibration-based weight compression. --- # GPTQ (Generative Pre-trained Transformer Quantization) Post-training quantization method that compresses LLMs to 4-bit with minimal accuracy loss using group-wise quantization. ## When to use GPTQ **Use GPTQ when:** - Need to fit large models (70B+) on limited GPU

26SKILL.mdUpdated Jul 9, 2026

frank-luongt/skills/codex/gptq

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/frank-luongt/faos-skills-marketplace.git

# Copy into Claude Code skills folder (global)
cp -r faos-skills-marketplace/skills/codex/databricks-mlflow-evaluation ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

frank-luongt/faos-skills-marketplace

26 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT