Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

eliferjunior/comet-ml

Name: comet-ml
Author: eliferjunior

.claude/skills/ts-comet-ml/SKILL.md

npx skillsauth add eliferjunior/Claude comet-ml

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Comet ML — ML Experiment Tracking & Model Management

Overview

Comet ML, the platform for tracking machine learning experiments, managing models, and monitoring production ML systems. Helps developers log experiments, compare model versions, and build reproducible ML pipelines with automatic code/data versioning.

Instructions

Experiment Tracking

# train.py — Track a training run with Comet ML
import comet_ml
from comet_ml import Experiment
import torch
from torch import nn, optim

# Initialize experiment — auto-logs git info, environment, and dependencies
experiment = Experiment(
    api_key=os.environ["COMET_API_KEY"],
    project_name="text-classifier",
    workspace="my-team",
    auto_metric_logging=True,
    auto_param_logging=True,
    auto_histogram_weight_logging=True,
    log_code=True,                          # Logs the full source code
)

# Log hyperparameters
experiment.log_parameters({
    "model": "distilbert-base-uncased",
    "learning_rate": 2e-5,
    "batch_size": 32,
    "epochs": 5,
    "optimizer": "AdamW",
    "warmup_steps": 100,
    "weight_decay": 0.01,
    "max_seq_length": 256,
})

# Log dataset information
experiment.log_dataset_hash(train_df)       # Hash for reproducibility
experiment.log_parameter("train_samples", len(train_df))
experiment.log_parameter("val_samples", len(val_df))

# Training loop with metric logging
for epoch in range(5):
    model.train()
    running_loss = 0

    for batch_idx, (inputs, labels) in enumerate(train_loader):
        optimizer.zero_grad()
        outputs = model(inputs)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()

        running_loss += loss.item()

        # Log batch-level metrics
        experiment.log_metric("batch_loss", loss.item(),
                             step=epoch * len(train_loader) + batch_idx)

    # Log epoch-level metrics
    avg_loss = running_loss / len(train_loader)
    experiment.log_metric("train_loss", avg_loss, epoch=epoch)

    # Validation
    val_loss, val_acc, val_f1 = evaluate(model, val_loader)
    experiment.log_metrics({
        "val_loss": val_loss,
        "val_accuracy": val_acc,
        "val_f1": val_f1,
    }, epoch=epoch)

    print(f"Epoch {epoch}: loss={avg_loss:.4f} val_acc={val_acc:.4f}")

# Log the model
experiment.log_model("text-classifier", "./model_weights.pt")

# Log confusion matrix
experiment.log_confusion_matrix(
    y_true=val_labels,
    y_predicted=val_predictions,
    labels=["negative", "neutral", "positive"],
)

# Log sample predictions
for i in range(10):
    experiment.log_text(f"Input: {val_texts[i]}\nPredicted: {val_predictions[i]}\nActual: {val_labels[i]}")

# Add tags for filtering in the dashboard
experiment.add_tags(["v2", "distilbert", "sentiment"])

experiment.end()

Model Registry

# Register and version models for deployment
from comet_ml import API

api = API(api_key=os.environ["COMET_API_KEY"])

# Register a model from an experiment
experiment = api.get_experiment("my-team", "text-classifier", "abc123")

# Register to model registry
experiment.register_model(
    model_name="text-classifier-prod",
    version="1.2.0",
    description="DistilBERT fine-tuned on customer reviews, F1=0.92",
    tags=["production-ready"],
    status="Production",              # Staging | Production | Archived
)

# Retrieve a registered model
model = api.get_model(workspace="my-team", model_name="text-classifier-prod")
latest = model.get_details("1.2.0")
print(f"Model: {latest['modelName']} v{latest['version']}")
print(f"Status: {latest['status']}")

# Download model assets
model.download("1.2.0", output_folder="./model_artifacts/")

# Compare two model versions
experiments = api.query("my-team", "text-classifier",
    Expression("val_f1", ">", 0.85))
for exp in experiments:
    print(f"{exp.id}: F1={exp.get_metric('val_f1')}")

LLM Tracking

# Track LLM/prompt experiments
from comet_ml import Experiment
from comet_ml.integration.openai import comet_openai_trace

# Auto-log OpenAI API calls
experiment = Experiment(project_name="llm-eval")
comet_openai_trace(experiment)              # Patches OpenAI client

# All OpenAI calls are automatically logged with:
# - Prompt text and system message
# - Model used and parameters
# - Response text
# - Token usage and cost
# - Latency

response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Summarize this text..."}],
)
# ↑ Automatically logged to Comet dashboard

# Manual prompt logging
experiment.log_text("prompt_template", "Summarize the following: {text}")
experiment.log_metric("prompt_tokens", response.usage.prompt_tokens)
experiment.log_metric("completion_tokens", response.usage.completion_tokens)
experiment.log_metric("total_cost", calculate_cost(response.usage))

Comparing Experiments

# Query and compare experiments programmatically
from comet_ml import API

api = API()

# Get all experiments in a project
experiments = api.get_experiments("my-team", "text-classifier")

# Filter by metric
best = sorted(experiments, key=lambda e: e.get_metric("val_f1") or 0, reverse=True)[:5]

for exp in best:
    params = exp.get_parameters_summary()
    metrics = exp.get_metrics_summary()
    print(f"Experiment: {exp.name}")
    print(f"  Model: {params.get('model')}")
    print(f"  F1: {metrics.get('val_f1')}")
    print(f"  Accuracy: {metrics.get('val_accuracy')}")

Installation

pip install comet-ml

# With framework integrations
pip install comet-ml[pytorch]
pip install comet-ml[tensorflow]
pip install comet-ml[sklearn]

Examples

Example 1: Setting up an evaluation pipeline for a RAG application

User request:

I have a RAG chatbot that answers questions from our docs. Set up Comet Ml to evaluate answer quality.

The agent creates an evaluation suite with appropriate metrics (faithfulness, relevance, answer correctness), configures test datasets from real user questions, runs baseline evaluations, and sets up CI integration so evaluations run on every prompt or retrieval change.

Example 2: Comparing model performance across prompts

User request:

We're testing GPT-4o vs Claude on our customer support prompts. Set up a comparison with Comet Ml.

The agent creates a structured experiment with the existing prompt set, configures both model providers, defines scoring criteria specific to customer support (accuracy, tone, completeness), runs the comparison, and generates a summary report with statistical significance indicators.

Guidelines

Log everything from the start — It's cheaper to log and not need it than to re-run experiments because you forgot a metric
Use auto-logging — Enable auto_metric_logging and auto_param_logging to capture framework metrics automatically
Tag experiments — Use tags (experiment.add_tags()) for filtering; tag by architecture, dataset version, or purpose
Model registry for deployment — Don't deploy from experiment artifacts; register models with versions and status tracking
Log dataset hashes — Use log_dataset_hash() for reproducibility; know exactly which data produced which results
Confusion matrices for classification — Always log confusion matrices; accuracy alone misses class imbalances
Compare before deploying — Use Comet's comparison view to verify new models outperform current production across all metrics
LLM cost tracking — Log token usage and calculated costs for every LLM call; costs add up fast in production

eliferjunior/comet-ml

.claude/skills/ts-comet-ml/SKILL.md

Expert guidance for Comet ML, the platform for tracking machine learning experiments, managing models, and monitoring production ML systems. Helps developers log experiments, compare model versions, and build reproducible ML pipelines with automatic code/data versioning.

development

Updated Apr 16, 2026

$ install --global

skillsauth

npx skillsauth add eliferjunior/Claude comet-ml

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 17, 2026, 1:27 AM6.6s1 file scanned

SKILL.md

name:: comet-ml
description:: Expert guidance for Comet ML, the platform for tracking machine learning experiments, managing models, and monitoring production ML systems. Helps developers log experiments, compare model versions, and build reproducible ML pipelines with automatic code/data versioning.
license:: Apache-2.0
compatibility:: No special requirements
author:: terminal-skills
version:: 1.0.0
category:: data-ai

Comet ML — ML Experiment Tracking & Model Management

Overview

Instructions

Experiment Tracking

# train.py — Track a training run with Comet ML
import comet_ml
from comet_ml import Experiment
import torch
from torch import nn, optim

# Initialize experiment — auto-logs git info, environment, and dependencies
experiment = Experiment(
    api_key=os.environ["COMET_API_KEY"],
    project_name="text-classifier",
    workspace="my-team",
    auto_metric_logging=True,
    auto_param_logging=True,
    auto_histogram_weight_logging=True,
    log_code=True,                          # Logs the full source code
)

# Log hyperparameters
experiment.log_parameters({
    "model": "distilbert-base-uncased",
    "learning_rate": 2e-5,
    "batch_size": 32,
    "epochs": 5,
    "optimizer": "AdamW",
    "warmup_steps": 100,
    "weight_decay": 0.01,
    "max_seq_length": 256,
})

# Log dataset information
experiment.log_dataset_hash(train_df)       # Hash for reproducibility
experiment.log_parameter("train_samples", len(train_df))
experiment.log_parameter("val_samples", len(val_df))

# Training loop with metric logging
for epoch in range(5):
    model.train()
    running_loss = 0

    for batch_idx, (inputs, labels) in enumerate(train_loader):
        optimizer.zero_grad()
        outputs = model(inputs)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()

        running_loss += loss.item()

        # Log batch-level metrics
        experiment.log_metric("batch_loss", loss.item(),
                             step=epoch * len(train_loader) + batch_idx)

    # Log epoch-level metrics
    avg_loss = running_loss / len(train_loader)
    experiment.log_metric("train_loss", avg_loss, epoch=epoch)

    # Validation
    val_loss, val_acc, val_f1 = evaluate(model, val_loader)
    experiment.log_metrics({
        "val_loss": val_loss,
        "val_accuracy": val_acc,
        "val_f1": val_f1,
    }, epoch=epoch)

    print(f"Epoch {epoch}: loss={avg_loss:.4f} val_acc={val_acc:.4f}")

# Log the model
experiment.log_model("text-classifier", "./model_weights.pt")

# Log confusion matrix
experiment.log_confusion_matrix(
    y_true=val_labels,
    y_predicted=val_predictions,
    labels=["negative", "neutral", "positive"],
)

# Log sample predictions
for i in range(10):
    experiment.log_text(f"Input: {val_texts[i]}\nPredicted: {val_predictions[i]}\nActual: {val_labels[i]}")

# Add tags for filtering in the dashboard
experiment.add_tags(["v2", "distilbert", "sentiment"])

experiment.end()

Model Registry

# Register and version models for deployment
from comet_ml import API

api = API(api_key=os.environ["COMET_API_KEY"])

# Register a model from an experiment
experiment = api.get_experiment("my-team", "text-classifier", "abc123")

# Register to model registry
experiment.register_model(
    model_name="text-classifier-prod",
    version="1.2.0",
    description="DistilBERT fine-tuned on customer reviews, F1=0.92",
    tags=["production-ready"],
    status="Production",              # Staging | Production | Archived
)

# Retrieve a registered model
model = api.get_model(workspace="my-team", model_name="text-classifier-prod")
latest = model.get_details("1.2.0")
print(f"Model: {latest['modelName']} v{latest['version']}")
print(f"Status: {latest['status']}")

# Download model assets
model.download("1.2.0", output_folder="./model_artifacts/")

# Compare two model versions
experiments = api.query("my-team", "text-classifier",
    Expression("val_f1", ">", 0.85))
for exp in experiments:
    print(f"{exp.id}: F1={exp.get_metric('val_f1')}")

LLM Tracking

# Track LLM/prompt experiments
from comet_ml import Experiment
from comet_ml.integration.openai import comet_openai_trace

# Auto-log OpenAI API calls
experiment = Experiment(project_name="llm-eval")
comet_openai_trace(experiment)              # Patches OpenAI client

# All OpenAI calls are automatically logged with:
# - Prompt text and system message
# - Model used and parameters
# - Response text
# - Token usage and cost
# - Latency

response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Summarize this text..."}],
)
# ↑ Automatically logged to Comet dashboard

# Manual prompt logging
experiment.log_text("prompt_template", "Summarize the following: {text}")
experiment.log_metric("prompt_tokens", response.usage.prompt_tokens)
experiment.log_metric("completion_tokens", response.usage.completion_tokens)
experiment.log_metric("total_cost", calculate_cost(response.usage))

Comparing Experiments

# Query and compare experiments programmatically
from comet_ml import API

api = API()

# Get all experiments in a project
experiments = api.get_experiments("my-team", "text-classifier")

# Filter by metric
best = sorted(experiments, key=lambda e: e.get_metric("val_f1") or 0, reverse=True)[:5]

for exp in best:
    params = exp.get_parameters_summary()
    metrics = exp.get_metrics_summary()
    print(f"Experiment: {exp.name}")
    print(f"  Model: {params.get('model')}")
    print(f"  F1: {metrics.get('val_f1')}")
    print(f"  Accuracy: {metrics.get('val_accuracy')}")

Installation

pip install comet-ml

# With framework integrations
pip install comet-ml[pytorch]
pip install comet-ml[tensorflow]
pip install comet-ml[sklearn]

Examples

Example 1: Setting up an evaluation pipeline for a RAG application

User request:

I have a RAG chatbot that answers questions from our docs. Set up Comet Ml to evaluate answer quality.

Example 2: Comparing model performance across prompts

User request:

We're testing GPT-4o vs Claude on our customer support prompts. Set up a comparison with Comet Ml.

Guidelines

Log everything from the start — It's cheaper to log and not need it than to re-run experiments because you forgot a metric
Use auto-logging — Enable auto_metric_logging and auto_param_logging to capture framework metrics automatically
Tag experiments — Use tags (experiment.add_tags()) for filtering; tag by architecture, dataset version, or purpose
Model registry for deployment — Don't deploy from experiment artifacts; register models with versions and status tracking
Log dataset hashes — Use log_dataset_hash() for reproducibility; know exactly which data produced which results
Confusion matrices for classification — Always log confusion matrices; accuracy alone misses class imbalances
Compare before deploying — Use Comet's comparison view to verify new models outperform current production across all metrics
LLM cost tracking — Log token usage and calculated costs for every LLM call; costs add up fast in production

Related Skills

eliferjunior/fireworks-ai

development

VerifiedTrustedCommunity

Expert guidance for Fireworks AI, the platform for running open-source LLMs (Llama, Mixtral, Qwen, etc.) with enterprise-grade speed and reliability. Helps developers integrate Fireworks' inference API, fine-tune models, and deploy custom model endpoints with function calling and structured output support.

SKILL.mdUpdated Apr 17, 2026

eliferjunior/fireworks-ai

eliferjunior/firecrawl

development

VerifiedTrustedCommunity

Convert any website into clean, structured data with Firecrawl — API-first web scraping service. Use when someone asks to "turn a website into markdown", "scrape website for LLM", "Firecrawl", "extract website content as clean text", "crawl and convert to structured data", or "scrape website for RAG". Covers single-page scraping, full-site crawling, structured extraction, and LLM-ready output.

SKILL.mdUpdated Apr 16, 2026

eliferjunior/firecrawl

eliferjunior/firebase

tools

VerifiedTrustedCommunity

Expert guidance for Firebase, Google's platform for building and scaling web and mobile applications. Helps developers set up authentication, Firestore/Realtime Database, Cloud Functions, hosting, storage, and analytics using Firebase's SDK and CLI.

SKILL.mdUpdated Apr 16, 2026

eliferjunior/firebase

eliferjunior/file-upload-processor

development

VerifiedTrustedCommunity

When the user needs to build file upload functionality for a web application. Use when the user mentions "file upload," "image upload," "upload endpoint," "multipart upload," "presigned URL," "S3 upload," "file validation," "upload to cloud storage," or "accept user files." Handles upload endpoints, file validation (type, size, magic bytes), cloud storage integration, and upload status tracking. For image/video processing after upload, see media-transcoder.

SKILL.mdUpdated Apr 16, 2026

eliferjunior/file-upload-processor

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/eliferjunior/Claude.git

# Copy into Claude Code skills folder (global)
cp -r Claude/.claude/skills/ts-comet-ml ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

eliferjunior/Claude

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT